الملخص
Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the visual dynamics induced by these actions. We introduce ViDAL, a Visual Dynamics-grounded Action Latent Space that anchors continuous action latents in the future visual dynamics of the scene. Specifically, ViDAL learns action latent space by training an Action Variational Autoencoder (Action VAE) to reconstruct action chunks while aligning its latent with future scene dynamics. When integrated into downstream robot policies, the proposed Action VAE serves as a plug-in action interface compatible with multiple VLA architectures and enables optional future-video prediction as an additional capability. Empirically, ViDAL outperforms competitive baselines on LIBERO with 98.1% average success, improves a multi-task $π_{0.5}$ policy on RoboTwin 2.0 from 54.3% to 65.5% (Clean) and from 33.2% to 43.1% (Random) success rates over 50 dual-arm tasks, and yields 20.0% and 23.4% absolute success-rate gains on real-world single-arm Franka and dual-arm ARX robot platforms.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Xu, Y., Chen, Y., Ma, Q., Yang, J., Li, P., Wang, K., Yang, J., Si, J., Huang, J., Liu, J., Liu, N., Huang, Y., & Wang, L. (2026). ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models. https://omanscience.com/ar/articles/vidal-a-visual-dynamics-grounded-action-latent-space-for-vision-language-action-models
MLA 9
Xu, Yuan, et al. "ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models." https://omanscience.com/ar/articles/vidal-a-visual-dynamics-grounded-action-latent-space-for-vision-language-action-models.
شيكاغو (المؤلف–التاريخ)
Xu, Yuan, Yixiang Chen, Qisen Ma, Jiabing Yang, Peiyan Li, Kai Wang, Jianhua Yang, Jianlou Si, Jun Huang, Jing Liu, Nianfeng Liu, Yan Huang, and Liang Wang. 2026. "ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models." https://omanscience.com/ar/articles/vidal-a-visual-dynamics-grounded-action-latent-space-for-vision-language-action-models.
هارفارد
Xu, Y., Chen, Y., Ma, Q., Yang, J., Li, P., Wang, K., Yang, J., Si, J., Huang, J., Liu, J., Liu, N., Huang, Y. and Wang, L. (2026) 'ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models', Available at: https://omanscience.com/ar/articles/vidal-a-visual-dynamics-grounded-action-latent-space-for-vision-language-action-models.
فانكوفر
Xu Y, Chen Y, Ma Q, Yang J, Li P, Wang K, et al. ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models. https://omanscience.com/ar/articles/vidal-a-visual-dynamics-grounded-action-latent-space-for-vision-language-action-models
IEEE
Y. Xu, Y. Chen, Q. Ma, J. Yang, P. Li, K. Wang, J. Yang, J. Si, J. Huang, J. Liu, N. Liu, Y. Huang, and L. Wang, "ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models," https://omanscience.com/ar/articles/vidal-a-visual-dynamics-grounded-action-latent-space-for-vision-language-action-models.