Abstract

Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-arts excel at long-horizon tasks, they struggle with the delicate and precise control required for complex tool manipulation. To overcome these limitations, we introduce P2P-T, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations. P2P-T bridges the cognitive and physical execution gap through a two-stage approach. First, pretraining an object-centric world model to extract stable pose priors; second, integrating these priors into an efficient, pose-aware low-level policy. By utilizing a robust automated data processing pipeline powered by modern foundation models, P2P-T completely bypasses the need for human-robot aligned data. This reduces overall training overhead drastically. With minimal per-task fine-tuning, our framework achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently remain out of reach for standard large-scale pretrained models.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wang, B., Wu, L., Wei, Y., Shao, S., Huang, C., Cui, W., Xu, Z., Wu, H., Chen, L., Ma, Y., & Li, H. (2026). From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations. https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations

MLA 9

Wang, Bangjun, et al. "From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations." https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations.

Chicago (author–date)

Wang, Bangjun, Longyan Wu, Yukun Wei, Shenghe Shao, Chaoyi Huang, Wenze Cui, Zetong Xu, Hanlin Wu, Long Chen, Yi Ma, and Hongyang Li. 2026. "From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations." https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations.

Harvard

Wang, B., Wu, L., Wei, Y., Shao, S., Huang, C., Cui, W., Xu, Z., Wu, H., Chen, L., Ma, Y. and Li, H. (2026) 'From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations', Available at: https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations.

Vancouver

Wang B, Wu L, Wei Y, Shao S, Huang C, Cui W, et al. From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations. https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations

IEEE

B. Wang, L. Wu, Y. Wei, S. Shao, C. Huang, W. Cui, Z. Xu, H. Wu, L. Chen, Y. Ma, and H. Li, "From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations," https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations.