Abstract

Shortcut learning is a prevalent issue in robot learning. The limited diversity of robot demonstration datasets can mislead policies into exploiting spurious correlations between tasks and irrelevant features, such as viewpoint or background. Collecting sufficiently diverse robot demonstrations is costly and inefficient, motivating algorithmic alternatives. We focus on vision-language-action (VLA) models and discover that different vision-language model backbones exhibit substantially different levels of susceptibility to visual shortcut learning. We find that model behavior correlates with our proposed representation-level metric, action margin, which requires no policy rollouts. Visual shortcuts consistently enter action representations in early layers, with models differing in the extent to which later layers correct them by incorporating language information. To boost models' attention to language, we introduce task scrubbing, a new domain-adversarial training method that decreases models' likelihood of using visual shortcuts and improves VLAs' generalization. Experiments in both simulation and the real world across multiple VLAs and visual cues show that task scrubbing improves out-of-distribution robustness and often eliminates visual shortcut learning.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Gerigk, J., Aspuru-Takata, K., Wu, C. H., Mohammadi, M., Zheng, S., & Gilitschenski, I. (2026). When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs. https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas

MLA 9

Gerigk, Jasper, et al. "When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs." https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas.

Chicago (author–date)

Gerigk, Jasper, Kenzo Aspuru-Takata, Chin-Hsuan Wu, Mohammad Mohammadi, Shuhong Zheng, and Igor Gilitschenski. 2026. "When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs." https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas.

Harvard

Gerigk, J., Aspuru-Takata, K., Wu, C. H., Mohammadi, M., Zheng, S. and Gilitschenski, I. (2026) 'When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs', Available at: https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas.

Vancouver

Gerigk J, Aspuru-Takata K, Wu CH, Mohammadi M, Zheng S, Gilitschenski I. When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs. https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas

IEEE

J. Gerigk, K. Aspuru-Takata, C. H. Wu, M. Mohammadi, S. Zheng, and I. Gilitschenski, "When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs," https://omanscience.com/en/articles/when-listening-becomes-easier-scrubbing-visual-cues-for-shortcut-free-vlas.