Abstract
Vision-language-action (VLA) policies infer grasp actions primarily from visual observations and robot state, but do not explicitly represent the physical response observed after contact. We present MIM-VLA, a motor-feedback-based architecture that encodes recent gripper current, position, velocity, and signal validity as a 128-dimensional interaction token. A motor-only Motor Interaction Module (MIM) is pretrained with human-reviewed contact and interaction-phase labels and then conditions only the gripper-action pathway of SmolVLA; arm actions and the position-control interface remain unchanged. The same token supports the MEM selector VLM that compares candidate interactions and produces evidence-conditioned selections and explanations. We evaluate MIM-VLA in three real-world settings: comparing the interaction resistance of visually different objects, disambiguating visually similar real and replica objects through active probing, and gently grasping fragile objects, including held-out instances. Across 13 object pairs, MIM-VLA selects the higher-resistance object in 75.0% of trials, compared with 48.8% for the SmolVLA baseline. For the evaluated tasks, the approach uses motor feedback already available from the gripper and does not require an additional tactile array, force-torque sensor, calibrated force estimate, or direct current control.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Lee, J., Koo, J., Kim, T., Cha, Y., & Choi, A. J. (2026). MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback. https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback
MLA 9
Lee, Jaeyoung, et al. "MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback." https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback.
Chicago (author–date)
Lee, Jaeyoung, Jiyeon Koo, Taehwa Kim, Yerin Cha, and Andrew Jaeyong Choi. 2026. "MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback." https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback.
Harvard
Lee, J., Koo, J., Kim, T., Cha, Y. and Choi, A. J. (2026) 'MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback', Available at: https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback.
Vancouver
Lee J, Koo J, Kim T, Cha Y, Choi AJ. MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback. https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback
IEEE
J. Lee, J. Koo, T. Kim, Y. Cha, and A. J. Choi, "MIM-VLA: Learning Physical Interaction Representations from Gripper Motor Feedback," https://omanscience.com/en/articles/mim-vla-learning-physical-interaction-representations-from-gripper-motor-feedback.