Abstract

Vision-language-action~(VLA) models have emerged as a promising paradigm for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a think with 2D first, drive with dedicated 3D priors paradigm. It first grounds 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support the trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning decisions, and construct the PlanningGrounding dataset to endow VLAs with planning-oriented grounding capability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of the explicit geometric chain-of-thought process for VLA-based planning.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Gui, X., Zhou, Y., Guo, D., Gong, J., Tan, F., & Shen, J. (2026). Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving. https://omanscience.com/en/articles/explicit-geometric-chain-of-thought-for-vision-language-action-in-autonomous-driving

MLA 9

Gui, Xingtai, et al. "Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving." https://omanscience.com/en/articles/explicit-geometric-chain-of-thought-for-vision-language-action-in-autonomous-driving.

Chicago (author–date)

Gui, Xingtai, Yucheng Zhou, Dongqian Guo, Jiahao Gong, Feiyang Tan, and Jianbing Shen. 2026. "Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving." https://omanscience.com/en/articles/explicit-geometric-chain-of-thought-for-vision-language-action-in-autonomous-driving.

Harvard

Gui, X., Zhou, Y., Guo, D., Gong, J., Tan, F. and Shen, J. (2026) 'Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving', Available at: https://omanscience.com/en/articles/explicit-geometric-chain-of-thought-for-vision-language-action-in-autonomous-driving.

Vancouver

Gui X, Zhou Y, Guo D, Gong J, Tan F, Shen J. Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving. https://omanscience.com/en/articles/explicit-geometric-chain-of-thought-for-vision-language-action-in-autonomous-driving

IEEE

X. Gui, Y. Zhou, D. Guo, J. Gong, F. Tan, and J. Shen, "Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving," https://omanscience.com/en/articles/explicit-geometric-chain-of-thought-for-vision-language-action-in-autonomous-driving.