Abstract

Vision-language-action (VLA) policies infer grasp actions primarily from visual observations and robot state, but do not explicitly represent the physical response observed after contact. We present MIM-VLA, a motor-feedback-based architecture that encodes recent gripper current, position, velocity, and signal validity as a 128-dimensional interaction token. A motor-only Motor Interaction Module (MIM) is pretrained with human-reviewed contact and interaction-phase labels and then conditions only the gripper-action pathway of SmolVLA; arm actions and the position-control interface remain unchanged. The same token supports the MEM selector VLM that compares candidate interactions and produces evidence-conditioned selections and explanations. We evaluate MIM-VLA in three real-world settings: comparing the interaction resistance of visually different objects, disambiguating visually similar real and replica objects through active probing, and gently grasping fragile objects, including held-out instances. Across 13 object pairs, MIM-VLA selects the higher-resistance object in 75.0% of trials, compared with 48.8% for the SmolVLA baseline. For the evaluated tasks, the approach uses motor feedback already available from the gripper and does not require an additional tactile array, force-torque sensor, calibrated force estimate, or direct current control.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Lee, J., Koo, J., Kim, T., Cha, Y., & Choi, A. J. (2026). MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback. https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback

MLA 9

Lee, Jaeyoung, et al. "MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback." https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback.

Chicago (author–date)

Lee, Jaeyoung, Jiyeon Koo, Taehwa Kim, Yerin Cha, and Andrew Jaeyong Choi. 2026. "MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback." https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback.

Harvard

Lee, J., Koo, J., Kim, T., Cha, Y. and Choi, A. J. (2026) 'MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback', Available at: https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback.

Vancouver

Lee J, Koo J, Kim T, Cha Y, Choi AJ. MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback. https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback

IEEE

J. Lee, J. Koo, T. Kim, Y. Cha, and A. J. Choi, "MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback," https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback.