الملخص
Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to capture fine-grained end-effector articulation, particularly finger-level geometric changes in human and dexterous robot hands. To address this limitation, we propose GALA, a Geometry-Aware Latent-Action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. However, naively incorporating point clouds yields fine-grained action representations with limited shared semantics, hindering cross-embodiment pretraining. To address this issue, we introduce the Unified End-effector Motion Representation (UEMR), which preserves fine-grained motion information while improving the cross-embodiment generalizability of latent actions. Building upon UEMR, GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos. Experiments on fine-grained motion probing, cross-embodiment retrieval, and downstream VLA evaluation demonstrate GALA's effectiveness in modeling generalizable fine-grained motions across embodiments, achieving 68.3% RoboCasa-GR1 success rate and 75.5% real-world success rate. Code, appendix, and demos are available at https://puzhenyuan.github.io/GALA-website/.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Liu, Y., Yuan, P., Zhu, X., Guo, Y., & Chen, J. (2026). GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments. https://omanscience.com/ar/articles/gala-geometry-aware-latent-action-modeling-for-vision-language-action-model-pretraining-across-embodiments
MLA 9
Liu, Yichen, et al. "GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments." https://omanscience.com/ar/articles/gala-geometry-aware-latent-action-modeling-for-vision-language-action-model-pretraining-across-embodiments.
شيكاغو (المؤلف–التاريخ)
Liu, Yichen, Puzhen Yuan, Xiang Zhu, Yanjiang Guo, and Jianyu Chen. 2026. "GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments." https://omanscience.com/ar/articles/gala-geometry-aware-latent-action-modeling-for-vision-language-action-model-pretraining-across-embodiments.
هارفارد
Liu, Y., Yuan, P., Zhu, X., Guo, Y. and Chen, J. (2026) 'GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments', Available at: https://omanscience.com/ar/articles/gala-geometry-aware-latent-action-modeling-for-vision-language-action-model-pretraining-across-embodiments.
فانكوفر
Liu Y, Yuan P, Zhu X, Guo Y, Chen J. GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments. https://omanscience.com/ar/articles/gala-geometry-aware-latent-action-modeling-for-vision-language-action-model-pretraining-across-embodiments
IEEE
Y. Liu, P. Yuan, X. Zhu, Y. Guo, and J. Chen, "GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments," https://omanscience.com/ar/articles/gala-geometry-aware-latent-action-modeling-for-vision-language-action-model-pretraining-across-embodiments.