الملخص

Humans typically rely on vision, touch, hearing, and proprioception to perceive contact and adapt their actions during manipulation. Providing robots with comparable responsiveness therefore requires hardware that can retain and use these complementary sensory signals. Most imitation-learning systems, however, observe demonstrations primarily through vision and proprioception, limiting access to contact information that is difficult to infer visually. We present PolyUMI, an open-source platform for scalable visual--tactile--audio demonstration collection and robot deployment. Its lightweight, wireless handheld gripper records synchronized wrist-camera, optical tactile, contact-audio, and proprioceptive observations without requiring a tethered workstation. The same sensing finger can be transferred to the robot end effector, preserving the sensing geometry between demonstration collection and policy execution. To effectively use these heterogeneous observations, we further introduce VisTA, a token-level multimodal policy that integrates information across sensors and time to predict contact-aware robot actions. Experiments spanning object inference, slip control, and contact-rich manipulation show that touch and audio reveal task-relevant information beyond vision and that VisTA is competitive with or outperforms existing multimodal policies. Together, PolyUMI and VisTA provide an accessible pipeline for collecting multimodal demonstrations and learning policies that perceive physical interaction beyond vision. Project Page: https://polyumi-vista.github.io

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Hayes, C. W., Krohn, R., Ramaswami, A., Ramaswami, A., Dengler, N., Lynch, K. M., Colgate, J. E., Chalvatzaki, G., & Elwin, M. L. (2026). PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation. https://omanscience.com/ar/articles/polyumi-accessible-visual-tactile-audio-data-collection-for-object-inference-and-manipulation

MLA 9

Hayes, Conor W., et al. "PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation." https://omanscience.com/ar/articles/polyumi-accessible-visual-tactile-audio-data-collection-for-object-inference-and-manipulation.

شيكاغو (المؤلف–التاريخ)

Hayes, Conor W., Rickmer Krohn, Aravind Ramaswami, Anunth Ramaswami, Nils Dengler, Kevin M. Lynch, J. Edward Colgate, Georgia Chalvatzaki, and Matthew L. Elwin. 2026. "PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation." https://omanscience.com/ar/articles/polyumi-accessible-visual-tactile-audio-data-collection-for-object-inference-and-manipulation.

هارفارد

Hayes, C. W., Krohn, R., Ramaswami, A., Ramaswami, A., Dengler, N., Lynch, K. M., Colgate, J. E., Chalvatzaki, G. and Elwin, M. L. (2026) 'PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation', Available at: https://omanscience.com/ar/articles/polyumi-accessible-visual-tactile-audio-data-collection-for-object-inference-and-manipulation.

فانكوفر

Hayes CW, Krohn R, Ramaswami A, Ramaswami A, Dengler N, Lynch KM, et al. PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation. https://omanscience.com/ar/articles/polyumi-accessible-visual-tactile-audio-data-collection-for-object-inference-and-manipulation

IEEE

C. W. Hayes, R. Krohn, A. Ramaswami, A. Ramaswami, N. Dengler, K. M. Lynch, J. E. Colgate, G. Chalvatzaki, and M. L. Elwin, "PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation," https://omanscience.com/ar/articles/polyumi-accessible-visual-tactile-audio-data-collection-for-object-inference-and-manipulation.