Abstract

Reusing manipulation experience across robot embodiments is important for scaling robot learning and reducing repeated task-specific data collection. However, changes in embodiment alter visual appearance, action dimensionality and semantics, and the whole-body configurations that can realize the same tool pose. We present SkelWAM, a skeleton-guided world-action model that couples perception and control through one explicit geometric representation for single-source cross-embodiment manipulation. Arm centerline geometry, tool-center-point (TCP) pose, and parallel-jaw commands form a shared 25-D state. The same definition underlies canonical third-person and wrist observations and future whole-body action targets. Trained with predictive visual supervision, a video-action mixture of transformers predicts canonical skeleton action chunks, which embodiment-specific constrained decoders convert into joint or continuum-robot controls. This formulation requires no one-to-one joint correspondence and uses no target-task demonstrations or target policy updates. We introduce LIBERO-Cross10, a source-only cross-embodiment transfer benchmark covering ten tasks and ten target embodiments across four morphological groups. On this benchmark, Franka-trained SkelWAM achieves 43.3% success over 1,000 episodes, exceeding the best-performing evaluated baseline by 36.2 percentage points. We further deploy a JAKA mini2-trained policy on the Feagine A03 continuum robot for three tabletop manipulation tasks, illustrating the approach's potential for real-world cross-embodiment manipulation. Project page: http://www.liukepku.com/skelwam/index.html

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Niu, P., Xie, Y., Peng, R., Zhao, H., & Liu, K. (2026). SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation. https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation

MLA 9

Niu, Pengjun, et al. "SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation." https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation.

Chicago (author–date)

Niu, Pengjun, Yujia Xie, Rui Peng, Hang Zhao, and Ke Liu. 2026. "SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation." https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation.

Harvard

Niu, P., Xie, Y., Peng, R., Zhao, H. and Liu, K. (2026) 'SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation', Available at: https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation.

Vancouver

Niu P, Xie Y, Peng R, Zhao H, Liu K. SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation. https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation

IEEE

P. Niu, Y. Xie, R. Peng, H. Zhao, and K. Liu, "SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation," https://omanscience.com/en/articles/skelwam-a-skeleton-guided-world-action-model-for-zero-shot-cross-embodiment-manipulation.