Abstract

Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While this fusion of language instructions and visual observations allows multimodal reasoning, it obscures how information is routed across modalities or what mechanisms drive navigation decisions. Thus, it remains unclear whether VLN models ground their predictions in relevant semantic cues or can track task progress. In this work, we study the interpretability and steerability of VLN models. We use intervention-based metrics that measure how visual observations, instructions, and visual memory causally influence navigation decisions. Our results show that these navigation policies are sensitive to all input modalities and do not depend on a single one. We further show that these agents encode navigation progress and retain semantic structure from their VLM backbones, enabling concept-level steering through internal activations. Finally, we extract activation vectors for abstract behaviors to transfer them zero-shot to out-of-distribution real-world scenarios, improving performance without additional fine-tuning.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Makowski, D. O., Gode, S., Nayak, A., Hutter, M., Schmid, C., Schmid, L. R., & Burgard, W. (2026). What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior. https://omanscience.com/en/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior

MLA 9

Makowski, Débora Oliveira, et al. "What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior." https://omanscience.com/en/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior.

Chicago (author–date)

Makowski, Débora Oliveira, Samiran Gode, Abhijeet Nayak, Marco Hutter, Cordelia Schmid, Lukas Rosenberger Schmid, and Wolfram Burgard. 2026. "What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior." https://omanscience.com/en/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior.

Harvard

Makowski, D. O., Gode, S., Nayak, A., Hutter, M., Schmid, C., Schmid, L. R. and Burgard, W. (2026) 'What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior', Available at: https://omanscience.com/en/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior.

Vancouver

Makowski DO, Gode S, Nayak A, Hutter M, Schmid C, Schmid LR, et al. What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior. https://omanscience.com/en/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior

IEEE

D. O. Makowski, S. Gode, A. Nayak, M. Hutter, C. Schmid, L. R. Schmid, and W. Burgard, "What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior," https://omanscience.com/en/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior.