الملخص
Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While this fusion of language instructions and visual observations allows multimodal reasoning, it obscures how information is routed across modalities or what mechanisms drive navigation decisions. Thus, it remains unclear whether VLN models ground their predictions in relevant semantic cues or can track task progress. In this work, we study the interpretability and steerability of VLN models. We use intervention-based metrics that measure how visual observations, instructions, and visual memory causally influence navigation decisions. Our results show that these navigation policies are sensitive to all input modalities and do not depend on a single one. We further show that these agents encode navigation progress and retain semantic structure from their VLM backbones, enabling concept-level steering through internal activations. Finally, we extract activation vectors for abstract behaviors to transfer them zero-shot to out-of-distribution real-world scenarios, improving performance without additional fine-tuning.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Makowski, D. O., Gode, S., Nayak, A., Hutter, M., Schmid, C., Schmid, L. R., & Burgard, W. (2026). What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior. https://omanscience.com/ar/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior
MLA 9
Makowski, Débora Oliveira, et al. "What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior." https://omanscience.com/ar/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior.
شيكاغو (المؤلف–التاريخ)
Makowski, Débora Oliveira, Samiran Gode, Abhijeet Nayak, Marco Hutter, Cordelia Schmid, Lukas Rosenberger Schmid, and Wolfram Burgard. 2026. "What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior." https://omanscience.com/ar/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior.
هارفارد
Makowski, D. O., Gode, S., Nayak, A., Hutter, M., Schmid, C., Schmid, L. R. and Burgard, W. (2026) 'What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior', Available at: https://omanscience.com/ar/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior.
فانكوفر
Makowski DO, Gode S, Nayak A, Hutter M, Schmid C, Schmid LR, et al. What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior. https://omanscience.com/ar/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior
IEEE
D. O. Makowski, S. Gode, A. Nayak, M. Hutter, C. Schmid, L. R. Schmid, and W. Burgard, "What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior," https://omanscience.com/ar/articles/what-do-vlm-based-vision-language-navigation-models-rely-on-interpreting-and-steering-policy-behavior.