الباحثون

Xiang Zhu

المنشورات 3

نسخة أولية وصول مفتوح

Video Prediction Policy 2: Predict Better, Act Better

Yanjiang Guo, Haodong Yan, Zhide Zhong وآخرون · 2026

World action models (WAMs) have emerged as an important class of generalist robot policies, aiming to transfer video prediction priors to action learning. However, we find that existing WAMs frequently produce incorrect motion predictions in open-ended environment, leading to erroneous actions. We attribute this limita …

نسخة أولية وصول مفتوح

GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments

Yichen Liu, Puzhen Yuan, Xiang Zhu وآخرون · 2026

Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to …

المؤلفون المشاركون