Abstract
Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-arts excel at long-horizon tasks, they struggle with the delicate and precise control required for complex tool manipulation. To overcome these limitations, we introduce P2P-T, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations. P2P-T bridges the cognitive and physical execution gap through a two-stage approach. First, pretraining an object-centric world model to extract stable pose priors; second, integrating these priors into an efficient, pose-aware low-level policy. By utilizing a robust automated data processing pipeline powered by modern foundation models, P2P-T completely bypasses the need for human-robot aligned data. This reduces overall training overhead drastically. With minimal per-task fine-tuning, our framework achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently remain out of reach for standard large-scale pretrained models.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Wang, B., Wu, L., Wei, Y., Shao, S., Huang, C., Cui, W., Xu, Z., Wu, H., Chen, L., Ma, Y., & Li, H. (2026). From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations. https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations
MLA 9
Wang, Bangjun, et al. "From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations." https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations.
Chicago (author–date)
Wang, Bangjun, Longyan Wu, Yukun Wei, Shenghe Shao, Chaoyi Huang, Wenze Cui, Zetong Xu, Hanlin Wu, Long Chen, Yi Ma, and Hongyang Li. 2026. "From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations." https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations.
Harvard
Wang, B., Wu, L., Wei, Y., Shao, S., Huang, C., Cui, W., Xu, Z., Wu, H., Chen, L., Ma, Y. and Li, H. (2026) 'From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations', Available at: https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations.
Vancouver
Wang B, Wu L, Wei Y, Shao S, Huang C, Cui W, et al. From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations. https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations
IEEE
B. Wang, L. Wu, Y. Wei, S. Shao, C. Huang, W. Cui, Z. Xu, H. Wu, L. Chen, Y. Ma, and H. Li, "From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations," https://omanscience.com/en/articles/from-pixel-to-poses-object-centric-tool-manipulation-learning-from-human-demonstrations.