Abstract
Robots that autonomously determine their actions from language instructions and sensory observations could perform new contact-rich manipulation tasks without task-specific training or hand-designed rules. To perform these tasks, robots must infer how objects contact one another and move as a result, then select actions. For contact inference and action selection, prior approaches involve designing estimation models and tactile feedback control laws, or learning models for object-motion estimation, action-outcome prediction, and action selection from tactile data. Instead, we propose TacZero, which uses a pretrained general-purpose vision-language model (VLM) to interpret visual and tactile observations and select robot actions without additional tactile or manipulation training or task-specific rules for contact interpretation or action selection. TacZero provides the VLM with camera images, robot state, and three-axis tactile responses represented as numerical values or vectors overlaid on the images. From these observations and interaction history, the VLM generates commands specifying target end-effector positions and gripper opening or closing, which a low-level controller executes. In real-world cylindrical-peg insertion experiments, TacZero succeeded in 15 of 20 trials with numerical tactile input, compared with 10 of 20 without tactile input. This study provides a concrete starting point for further research on contact-rich manipulation using general-purpose VLMs and highlights challenges in pursuing this direction.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Tanaka, K. (2026). TacZero: Training-Free Peg Insertion Using a General-Purpose Vision-Language Model with Tactile Feedback. https://omanscience.com/en/articles/taczero-training-free-peg-insertion-using-a-general-purpose-vision-language-model-with-tactile-feedback
MLA 9
Tanaka, Kazutoshi. "TacZero: Training-Free Peg Insertion Using a General-Purpose Vision-Language Model with Tactile Feedback." https://omanscience.com/en/articles/taczero-training-free-peg-insertion-using-a-general-purpose-vision-language-model-with-tactile-feedback.
Chicago (author–date)
Tanaka, Kazutoshi. 2026. "TacZero: Training-Free Peg Insertion Using a General-Purpose Vision-Language Model with Tactile Feedback." https://omanscience.com/en/articles/taczero-training-free-peg-insertion-using-a-general-purpose-vision-language-model-with-tactile-feedback.
Harvard
Tanaka, K. (2026) 'TacZero: Training-Free Peg Insertion Using a General-Purpose Vision-Language Model with Tactile Feedback', Available at: https://omanscience.com/en/articles/taczero-training-free-peg-insertion-using-a-general-purpose-vision-language-model-with-tactile-feedback.
Vancouver
Tanaka K. TacZero: Training-Free Peg Insertion Using a General-Purpose Vision-Language Model with Tactile Feedback. https://omanscience.com/en/articles/taczero-training-free-peg-insertion-using-a-general-purpose-vision-language-model-with-tactile-feedback
IEEE
K. Tanaka, "TacZero: Training-Free Peg Insertion Using a General-Purpose Vision-Language Model with Tactile Feedback," https://omanscience.com/en/articles/taczero-training-free-peg-insertion-using-a-general-purpose-vision-language-model-with-tactile-feedback.