الباحثون

Samiran Gode

المنشورات 1

نسخة أولية وصول مفتوح

What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior

Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While this fusion of language instructions and visual observations allows multimodal reasoning, it obscures how information is routed across modalities or what mechanisms drive na …

المؤلفون المشاركون