الباحثون

Huchuan Lu

المنشورات 8

نسخة أولية وصول مفتوح

Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models

Zaibin Zhang, Binghao Ran, Yuhan Wu وآخرون · 2026

Generalization in multi-arm collaboration can be studied as composing familiar atomic skills in new ways across arms. However, existing evaluations offer limited insight into which training and architectural choices support this ability under different coordination requirements. We introduce \textbf{ACG-Bench}, a bench …

نسخة أولية وصول مفتوح

When to Retrieve, When to Stay: Uncertainty-Aware Temporal Evidence Allocation for Streaming Video-LLMs

Xiang Hu, Jiazuo Yu, Lu Zhang وآخرون · 2026

Streaming video understanding requires Video Large Language Models (Video-LLMs) to reason over continuous visual streams under causal constraints. As the visual history grows, a bounded visual?processing budget requires evidence selection that balances temporal recency with query relevance. Recent-only selection exclud …

نسخة أولية وصول مفتوح

WorldGuide: Learning Success-Failure Boundaries in Latent World Models for Vision-Language-Action Policies

Lin Liu, Lu Zhang, Ziying Song وآخرون · 2026

Latent world models offer a promising way to improve Vision-Language-Action policies by capturing the consequences of actions. However, models trained primarily on expert demonstrations have limited exposure to failure outcomes and may struggle to distinguish visually similar successful and failed interactions. We prop …

نسخة أولية وصول مفتوح

UCON: Uncertainty-aware Navigation with Historical Re-association in Dynamic Environments

Bing Sun, Yue Lin, Yongsheng Yuan وآخرون · 2026

Autonomous navigation in dynamic environments is hindered by two fundamental challenges: perception instability and uncertainty-optimization mismatch. The former leads to identity switches and unreliable motion estimation, while the latter prevents principled incorporation of motion uncertainty into trajectory optimiza …

نسخة أولية وصول مفتوح

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies

Junlan Xiao, Junwei Jiang, Zaibin Zhang وآخرون · 2026

Vision-Language-Action (VLA) policies achieve strong performance in robotic manipulation but remain brittle once execution deviates from nominal trajectories. We propose CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution. Instead of gene …

نسخة أولية وصول مفتوح

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

Zhongbo Zhang, Jiayi Jin, Yifan Wang وآخرون · 2026

Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedur …

نسخة أولية وصول مفتوح

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

Zhongbo Zhang, Zaibin Zhang, Yifan Wang وآخرون · 2026

3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasible now, but also anticipating where the interaction is heading. Existing policies largely leave such foresight to emerge implicitly from act …

المؤلفون المشاركون