Abstract

Terrain assessment is a critical capability for off-road mobile robots, enabling safe and reliable navigation through unstructured and geometrically complex environments. Conventional geometry-based terrain assessment is fast to compute but often overly conservative in unstructured environments. We present PIVOT: a Physically Informed Vision-Language Off-Road Traversability navigation system that augments conventional geometry-based planning with vision-language-model (VLM)-based semantic reasoning for field robots. To physically ground this assessment, we quantify how strongly the VLM's predicted traversal energy cost, robot vibration, and wheel slip correlate with real-world measurements and introduce a unified traversability score that weights each modality by its prediction-measurement correlation. For efficiency, we design a two-level navigation architecture that retains geometry-based planning as the nominal mode and invokes semantic replanning only when that mode fails to find a path. Across five repeated closed-loop trials on a mixed-terrain route totalling around $6.4$ km, the proposed system increases overall autonomy from $59.6\%$ to $97.0\%$, reduces human interventions from $11$ to $3$, and increases the mean distance between interventions (MDBI) from $69.2$ m to $412.9$ m compared with geometry-only navigation. These results demonstrate that physically grounded VLM-based terrain assessment can substantially extend autonomous navigation beyond the limitations of geometry alone, while preserving efficient geometric planning as the nominal mode.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Jiao, A., Zhao, W., Sahak, H., & Barfoot, T. D. (2026). PIVOT: Physically Informed Vision-Language Off-Road Traversability for Field Robot Navigation. https://omanscience.com/en/articles/pivot-physically-informed-vision-language-off-road-traversability-for-field-robot-navigation

MLA 9

Jiao, Aoran, et al. "PIVOT: Physically Informed Vision-Language Off-Road Traversability for Field Robot Navigation." https://omanscience.com/en/articles/pivot-physically-informed-vision-language-off-road-traversability-for-field-robot-navigation.

Chicago (author–date)

Jiao, Aoran, Wenda Zhao, Hshmat Sahak, and Timothy D. Barfoot. 2026. "PIVOT: Physically Informed Vision-Language Off-Road Traversability for Field Robot Navigation." https://omanscience.com/en/articles/pivot-physically-informed-vision-language-off-road-traversability-for-field-robot-navigation.

Harvard

Jiao, A., Zhao, W., Sahak, H. and Barfoot, T. D. (2026) 'PIVOT: Physically Informed Vision-Language Off-Road Traversability for Field Robot Navigation', Available at: https://omanscience.com/en/articles/pivot-physically-informed-vision-language-off-road-traversability-for-field-robot-navigation.

Vancouver

Jiao A, Zhao W, Sahak H, Barfoot TD. PIVOT: Physically Informed Vision-Language Off-Road Traversability for Field Robot Navigation. https://omanscience.com/en/articles/pivot-physically-informed-vision-language-off-road-traversability-for-field-robot-navigation

IEEE

A. Jiao, W. Zhao, H. Sahak, and T. D. Barfoot, "PIVOT: Physically Informed Vision-Language Off-Road Traversability for Field Robot Navigation," https://omanscience.com/en/articles/pivot-physically-informed-vision-language-off-road-traversability-for-field-robot-navigation.