الملخص
Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned using language, points, or box prompts, which are first transformed into a shared object-centric representation of the selected objects. This representation captures target-focused visual features and metric object geometry, which is mixed with robot state history through a multi-stream transformer to predict action chunks. Although GAMs can be run autonomously, they can also serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation. On RoboTwin 2.0, GAM achieves an average success rate of 55.3% across 50 tasks (vs. 52.0% for Spatial Forcing), including 47.6% under scene randomization (vs. 30.4% for Abot-M0), with its action policy trained only on clean-scene demonstrations. On LIBERO-PRO, it achieves a state-of-the-art average success rate of 61% (vs. 53% for $π_{0.5}$) across 16 perturbation settings, with the largest gains when targets are relocated or newly designated. On two real robots, GAM retains 17/20 successes under visual shift on a bimanual YAM versus 4/20 for $π_{0.5}$, while its composition with a Molmo2 planner on a Franka achieves 64.7% ID and 49.8% OOD step completion on long-horizon and memory-dependent tasks.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Zhang, G., Huang, W., Shailesh, S., Peng, Y., Duan, J., & Krishna, R. (2026). Grounded Action Model: 3D Grounding as a Foundation for Robotics. https://omanscience.com/ar/articles/grounded-action-model-3d-grounding-as-a-foundation-for-robotics
MLA 9
Zhang, Gehao, et al. "Grounded Action Model: 3D Grounding as a Foundation for Robotics." https://omanscience.com/ar/articles/grounded-action-model-3d-grounding-as-a-foundation-for-robotics.
شيكاغو (المؤلف–التاريخ)
Zhang, Gehao, Weikai Huang, Shailesh Shailesh, Yiyan Peng, Jiafei Duan, and Ranjay Krishna. 2026. "Grounded Action Model: 3D Grounding as a Foundation for Robotics." https://omanscience.com/ar/articles/grounded-action-model-3d-grounding-as-a-foundation-for-robotics.
هارفارد
Zhang, G., Huang, W., Shailesh, S., Peng, Y., Duan, J. and Krishna, R. (2026) 'Grounded Action Model: 3D Grounding as a Foundation for Robotics', Available at: https://omanscience.com/ar/articles/grounded-action-model-3d-grounding-as-a-foundation-for-robotics.
فانكوفر
Zhang G, Huang W, Shailesh S, Peng Y, Duan J, Krishna R. Grounded Action Model: 3D Grounding as a Foundation for Robotics. https://omanscience.com/ar/articles/grounded-action-model-3d-grounding-as-a-foundation-for-robotics
IEEE
G. Zhang, W. Huang, S. Shailesh, Y. Peng, J. Duan, and R. Krishna, "Grounded Action Model: 3D Grounding as a Foundation for Robotics," https://omanscience.com/ar/articles/grounded-action-model-3d-grounding-as-a-foundation-for-robotics.