Abstract

Zero-shot vision-and-language navigation in continuous environments (VLN-CE) increasingly places foundation vision-language models (VLMs) inside the navigation loop. Existing systems commonly request cardinal outputs such as waypoints, pixels, headings, progress values, or absolute arrival decisions, coupling a generative response to geometric magnitude or an irreversible commitment. We study a complementary model-robot interface: the VLM compares controller-constructed alternatives, while geometry, thresholds, action magnitude, and execution remain on the physical side. We instantiate this idea in C2Nav, a training-free framework with three coordinated faculties. Seeing performs ordinal Gaze Election over physically vetted candidate views; Remembering maintains a compact route sketch and compares adjacent instruction-leg hypotheses; and Arriving combines a hesitation ladder, look-back comparison, and revocable walk-back for reliable stopping. On the public OpenNav R2R-CE 100 protocol, C2Nav with Qwen3-VL-8B-Instruct obtains 41.0% OSR, 31.0% SR, and 16.7% SPL, while the same interface with the standard GPT-5.5 model reaches 54.0% OSR, 44.0% SR, and 29.0% SPL. Whole-faculty ablations reduce SR to 14.0% without Seeing, 25.0% without Remembering, and 29.0% without Arriving. Matched role inversions that replace only the comparative answer form with cardinal/absolute questions reduce SR to 12.0%, 28.0%, and 21.0% in the spatial, transition, and terminal slots, respectively. The results indicate that a constrained decision interface and stronger VLM reasoning are complementary rather than interchangeable.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zheng, R., Zhang, C., & Liu, Y. (2026). C$^2$Nav: Compare Before You Commit for Zero-Shot Vision-and-Language Navigation. https://omanscience.com/en/articles/c-2-nav-compare-before-you-commit-for-zero-shot-vision-and-language-navigation

MLA 9

Zheng, Runtian, et al. "C$^2$Nav: Compare Before You Commit for Zero-Shot Vision-and-Language Navigation." https://omanscience.com/en/articles/c-2-nav-compare-before-you-commit-for-zero-shot-vision-and-language-navigation.

Chicago (author–date)

Zheng, Runtian, Congpeng Zhang, and Ying Liu. 2026. "C$^2$Nav: Compare Before You Commit for Zero-Shot Vision-and-Language Navigation." https://omanscience.com/en/articles/c-2-nav-compare-before-you-commit-for-zero-shot-vision-and-language-navigation.

Harvard

Zheng, R., Zhang, C. and Liu, Y. (2026) 'C$^2$Nav: Compare Before You Commit for Zero-Shot Vision-and-Language Navigation', Available at: https://omanscience.com/en/articles/c-2-nav-compare-before-you-commit-for-zero-shot-vision-and-language-navigation.

Vancouver

Zheng R, Zhang C, Liu Y. C$^2$Nav: Compare Before You Commit for Zero-Shot Vision-and-Language Navigation. https://omanscience.com/en/articles/c-2-nav-compare-before-you-commit-for-zero-shot-vision-and-language-navigation

IEEE

R. Zheng, C. Zhang, and Y. Liu, "C$^2$Nav: Compare Before You Commit for Zero-Shot Vision-and-Language Navigation," https://omanscience.com/en/articles/c-2-nav-compare-before-you-commit-for-zero-shot-vision-and-language-navigation.