الباحثون

Xiang Zhang

المنشورات 9

نسخة أولية وصول مفتوح

Looking Inside LLMs: Small-World Connectivity as a Signature of Reasoning Performance

Zheng Huang, Sansheng Cao, Enpei Zhang وآخرون · 2026

Understanding large language model (LLM) reasoning requires looking beyond behavioral performance to examine how reasoning ability is reflected in internal organization. Inspired by neuroscience findings linking higher intelligence to stronger small-world organization in functional brain networks, we investigate small- …

نسخة أولية وصول مفتوح

OverLay++: Dense-Overlap Layout-to-Image Generation Dataset

Layout-to-Image generation has made substantial progress in spatial and object-level control. However, existing methods still struggle with complex scenes containing many overlapping and interacting objects. We argue that training data is a particular bottleneck: existing datasets lack examples with dense, complex obje …

نسخة أولية وصول مفتوح

MemTrace: State-Consistent Memory for Long-Horizon Coding Agents

Hongming Xu, Le Zhou, ZhongHe Jin وآخرون · 2026

As coding agents take on long-horizon software evolution tasks spanning multiple files and stages, longer execution trajectories introduce two coupled challenges: (1) accumulated histories strain context budgets, and (2) repository changes can invalidate earlier execution evidence. Existing approaches address these cha …

نسخة أولية وصول مفتوح

CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation

Hanwen Lu, Jun He, Mingjia Yang وآخرون · 2026

Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR-his, a decade-spanning cross-view dataset containing 43,653 …

نسخة أولية وصول مفتوح

LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

Shengxiang Ji, Boyang Wang, Haiyang Xu وآخرون · 2026

We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often …

نسخة أولية وصول مفتوح

Affordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language Navigation

Xuekang Yang, Lu Chen, Shuang Luo وآخرون · 2026

Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow passage does not ensure traversal; the agent must identify …

نسخة أولية وصول مفتوح

What Can a Leaderboard Certify? Compositional Controllability for Fair Evaluation and Training of Biomedical Literature-Review Agents

Zhaowei Han, Xiang Zhang, Lingxiao Guan وآخرون · 2026

Leaderboards rank long-horizon agents by their final outputs. Yet a higher score alone does not establish whether two systems are comparable or which stage accounts for the difference. Unequal evidence, inputs, or budgets can affect scores, and statistical corrections do not remove this mismatch. We introduce compositi …

المؤلفون المشاركون