الباحثون

شان يو

المنشورات 2

نسخة أولية وصول مفتوح

DSDyn-VLA: A Dual-Stream Dynamic Manipulation Framework with Motion Perception, Future Awareness, and Realtime Correction

وينهاو لي, Xiu Su, يو هان وآخرون · 2026

While Vision-Language-Action (VLA) models excel in static tasks, they struggle in dynamic environments where objects are in motion (e.g., conveyor belt manipulation). We identify three fundamental limitations hindering current VLAs in these scenarios: the \textbf{perception gap}, where static visual inputs lack tempora …

نسخة أولية وصول مفتوح

IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models

Open robot policies increasingly follow two paradigms: vision-language-action models (VLAs) directly map observations and instructions to actions, whereas world-action models (WAMs) incorporate learned video or world dynamics into policy learning or action generation. Although both target the same manipulation tasks an …

المؤلفون المشاركون