Abstract
Reusing manipulation experience across robot embodiments is important for scaling robot learning and reducing repeated task-specific data collection. However, changes in embodiment alter visual appearance, action dimensionality and semantics, and the whole-body configurations that can realize the same tool pose. We present SkelWAM, a skeleton-guided world-action model that couples perception and control through one explicit geometric representation for single-source cross-embodiment manipulation. Arm centerline geometry, tool-center-point (TCP) pose, and parallel-jaw commands form a shared 25-D state. The same definition underlies canonical third-person and wrist observations and future whole-body action targets. Trained with predictive visual supervision, a video-action mixture of transformers predicts canonical skeleton action chunks, which embodiment-specific constrained decoders convert into joint or continuum-robot controls. This formulation requires no one-to-one joint correspondence and uses no target-task demonstrations or target policy updates. We introduce LIBERO-Cross10, a source-only cross-embodiment transfer benchmark covering ten tasks and ten target embodiments across four morphological groups. On this benchmark, Franka-trained SkelWAM achieves 43.3% success over 1,000 episodes, exceeding the best-performing evaluated baseline by 36.2 percentage points. We further deploy a JAKA mini2-trained policy on the Feagine A03 continuum robot for three tabletop manipulation tasks, illustrating the approach's potential for real-world cross-embodiment manipulation. Project page: http://www.liukepku.com/skelwam/index.html
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Niu, P., Xie, Y., Peng, R., Zhao, H., & Liu, K. (2026). SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation. https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation
MLA 9
Niu, Pengjun, et al. "SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation." https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation.
Chicago (author–date)
Niu, Pengjun, Yujia Xie, Rui Peng, Hang Zhao, and Ke Liu. 2026. "SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation." https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation.
Harvard
Niu, P., Xie, Y., Peng, R., Zhao, H. and Liu, K. (2026) 'SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation', Available at: https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation.
Vancouver
Niu P, Xie Y, Peng R, Zhao H, Liu K. SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation. https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation
IEEE
P. Niu, Y. Xie, R. Peng, H. Zhao, and K. Liu, "SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation," https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation.