الباحثون

Xiaojuan Qi

المنشورات 6

نسخة أولية وصول مفتوح

Long-WAM: Scaling the Context of World-Action Models

Wei Huang, Bohan Zhang, Chenzhi Liu وآخرون · 2026

Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is …

نسخة أولية وصول مفتوح

MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation

Chenzhi Liu, Yue Zhang, Jiehong Lin وآخرون · 2026

Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermor …

نسخة أولية وصول مفتوح

D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation

Zijian Ye, Chengqi Wei, Wei Huang وآخرون · 2026

Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action (VLA) policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model (VLM) pass. We present D$^ …

نسخة أولية وصول مفتوح

AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models

Rui Wang, Xiangyu Wang, Donglin Yang وآخرون · 2026

World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions a …

المؤلفون المشاركون