الملخص
Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated by a simple premise: hands can change, but the underlying physics of interaction does not. We leverage human tactile data to improve dexterous manipulation policies. Specifically, we first build a tactile motion-capture system that synchronously records images, tactile signals, and hand motions. Using this system, we construct the UVTA dataset spanning five contact-rich tasks, with 1,000 human demonstrations covering diverse interaction patterns and 150 robot demonstrations per task. To transfer the underlying physics of human interaction to robot control, we propose a Unified Visual-Tactile-Action Model that maps both embodiments into aligned tactile and action representations and jointly predicts future action and tactile trajectories. The joint objective enables human demonstrations to supervise contact-aware representation learning, while only robot actions are executed during deployment. In real-robot evaluations across five tasks, our method achieves an average success rate of 70%, outperforming the strongest visual-tactile baseline, which achieves 29%, and an architecture ablation, which achieves 42%. Performance improves consistently with additional human demonstrations and exhibits no saturation at 1,000 demonstrations per task, validating the effectiveness of scalable human tactile data for dexterous manipulation. Project page is available at https://uni-vta.github.io/.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Li, W., Zhao, Q., Hao, J., Zhu, X., Liu, T., Zhang, K., Wen, C., & Huang, S. (2026). Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation. https://omanscience.com/ar/articles/unified-visual-tactile-action-modeling-from-human-demonstrations-for-dexterous-manipulation
MLA 9
Li, Wenqiao, et al. "Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation." https://omanscience.com/ar/articles/unified-visual-tactile-action-modeling-from-human-demonstrations-for-dexterous-manipulation.
شيكاغو (المؤلف–التاريخ)
Li, Wenqiao, Qianyou Zhao, Jiawen Hao, Xuezhou Zhu, Tengyu Liu, Kaifeng Zhang, Chuan Wen, and Siyuan Huang. 2026. "Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation." https://omanscience.com/ar/articles/unified-visual-tactile-action-modeling-from-human-demonstrations-for-dexterous-manipulation.
هارفارد
Li, W., Zhao, Q., Hao, J., Zhu, X., Liu, T., Zhang, K., Wen, C. and Huang, S. (2026) 'Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation', Available at: https://omanscience.com/ar/articles/unified-visual-tactile-action-modeling-from-human-demonstrations-for-dexterous-manipulation.
فانكوفر
Li W, Zhao Q, Hao J, Zhu X, Liu T, Zhang K, et al. Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation. https://omanscience.com/ar/articles/unified-visual-tactile-action-modeling-from-human-demonstrations-for-dexterous-manipulation
IEEE
W. Li, Q. Zhao, J. Hao, X. Zhu, T. Liu, K. Zhang, C. Wen, and S. Huang, "Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation," https://omanscience.com/ar/articles/unified-visual-tactile-action-modeling-from-human-demonstrations-for-dexterous-manipulation.