Abstract
Shortcut learning is a prevalent issue in robot learning. The limited diversity of robot demonstration datasets can mislead policies into exploiting spurious correlations between tasks and irrelevant features, such as viewpoint or background. Collecting sufficiently diverse robot demonstrations is costly and inefficient, motivating algorithmic alternatives. We focus on vision-language-action (VLA) models and discover that different vision-language model backbones exhibit substantially different levels of susceptibility to visual shortcut learning. We find that model behavior correlates with our proposed representation-level metric, action margin, which requires no policy rollouts. Visual shortcuts consistently enter action representations in early layers, with models differing in the extent to which later layers correct them by incorporating language information. To boost models' attention to language, we introduce task scrubbing, a new domain-adversarial training method that decreases models' likelihood of using visual shortcuts and improves VLAs' generalization. Experiments in both simulation and the real world across multiple VLAs and visual cues show that task scrubbing improves out-of-distribution robustness and often eliminates visual shortcut learning.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Gerigk, J., Aspuru-Takata, K., Wu, C. H., Mohammadi, M., Zheng, S., & Gilitschenski, I. (2026). When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs. https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas
MLA 9
Gerigk, Jasper, et al. "When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs." https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas.
Chicago (author–date)
Gerigk, Jasper, Kenzo Aspuru-Takata, Chin-Hsuan Wu, Mohammad Mohammadi, Shuhong Zheng, and Igor Gilitschenski. 2026. "When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs." https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas.
Harvard
Gerigk, J., Aspuru-Takata, K., Wu, C. H., Mohammadi, M., Zheng, S. and Gilitschenski, I. (2026) 'When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs', Available at: https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas.
Vancouver
Gerigk J, Aspuru-Takata K, Wu CH, Mohammadi M, Zheng S, Gilitschenski I. When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs. https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas
IEEE
J. Gerigk, K. Aspuru-Takata, C. H. Wu, M. Mohammadi, S. Zheng, and I. Gilitschenski, "When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs," https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas.