الملخص
The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This unified action space improves the world model's performance when adapted to previously unseen robot embodiments. We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels. Our results show that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting to new robots. We evaluate both models on dexterous manipulation tasks and a modified LIBERO benchmark. LAC-WM improves downstream performance over EAC-WM by up to 46.7% on dexterous manipulation and 11.7% on LIBERO. Crucially, the unified latent action space allows LAC-WM's downstream performance to scale positively with the number of embodiments used during pretraining. In contrast, the disjoint action space in EAC-WM leads to decreased performance as the number of pretraining embodiments increases. These results highlight the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics. Project website: https://lacwm.github.io/
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Huang, H., Yenamandra, S., Majumdar, A., Aljalbout, E., Nagarajan, T., Yang, T. Y., Rai, A., Rabbat, M., Fei-Fei, L., Wu, J., Wu, T., & Meier, F. (2026). Cross-Embodiment Robot Foundation World Models with Latent Actions. https://omanscience.com/ar/articles/cross-embodiment-robot-foundation-world-models-with-latent-actions
MLA 9
Huang, Huang, et al. "Cross-Embodiment Robot Foundation World Models with Latent Actions." https://omanscience.com/ar/articles/cross-embodiment-robot-foundation-world-models-with-latent-actions.
شيكاغو (المؤلف–التاريخ)
Huang, Huang, Sriram Yenamandra, Arjun Majumdar, Elie Aljalbout, Tushar Nagarajan, Tsung-Yen Yang, Akshara Rai, Michael Rabbat, Li Fei-Fei, Jiajun Wu, Tingfan Wu, and Franziska Meier. 2026. "Cross-Embodiment Robot Foundation World Models with Latent Actions." https://omanscience.com/ar/articles/cross-embodiment-robot-foundation-world-models-with-latent-actions.
هارفارد
Huang, H., Yenamandra, S., Majumdar, A., Aljalbout, E., Nagarajan, T., Yang, T. Y., Rai, A., Rabbat, M., Fei-Fei, L., Wu, J., Wu, T. and Meier, F. (2026) 'Cross-Embodiment Robot Foundation World Models with Latent Actions', Available at: https://omanscience.com/ar/articles/cross-embodiment-robot-foundation-world-models-with-latent-actions.
فانكوفر
Huang H, Yenamandra S, Majumdar A, Aljalbout E, Nagarajan T, Yang TY, et al. Cross-Embodiment Robot Foundation World Models with Latent Actions. https://omanscience.com/ar/articles/cross-embodiment-robot-foundation-world-models-with-latent-actions
IEEE
H. Huang, S. Yenamandra, A. Majumdar, E. Aljalbout, T. Nagarajan, T. Y. Yang, A. Rai, M. Rabbat, L. Fei-Fei, J. Wu, T. Wu, and F. Meier, "Cross-Embodiment Robot Foundation World Models with Latent Actions," https://omanscience.com/ar/articles/cross-embodiment-robot-foundation-world-models-with-latent-actions.