الملخص
Driving vision-language-action (VLA) models increasingly reason before acting, but their intermediate reasoning is often weakly grounded in physical scene evidence and loosely connected to executable behavior. We present GRAVA, a framework built around Grounded Reasoning-to-Action (GRA), which unifies grounding, reasoning, and action generation in a single autoregressive stream. GRA links action-relevant language references to 2D visual regions and ego-centric physical states, organizes object interactions and decisions in a trajectory-anchored typed graph, and serializes this structure into grounded reasoning. A single VLM generates this reasoning followed by a compact Executable Planner action that is deterministically decoded into a continuous trajectory. We further introduce an agentic GRA data construction pipeline that combines forward scene grounding with backward trajectory anchoring, and use it to build GR-NavSim with 2.2M grounded question-answer pairs and 70K GRA reasoning traces. A progressive training strategy develops grounded cognition through pre-training, establishes the reasoning-to-action interface through imitation, and improves driving behavior through reinforcement learning and exploration. Using about 60% of the available human driving demonstrations for action supervision, GRAVA-8B achieves state-of-the-art performance among purely autoregressive driving models on the full NAVSIM benchmark. On an internal long-tail benchmark, full GRA improves key-object compliance and Closed-loop Driving Score by 19.3% and 20.5% over action-only prediction, respectively. These results show the benefit of preserving action-relevant physical evidence from grounded reasoning through executable action generation.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Liu, X., Li, H., Leng, J., Wang, L., & Sun, C. (2026). GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving. https://omanscience.com/ar/articles/grava-grounded-reasoning-to-action-representation-and-learning-for-autonomous-driving
MLA 9
Liu, Xiao, et al. "GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving." https://omanscience.com/ar/articles/grava-grounded-reasoning-to-action-representation-and-learning-for-autonomous-driving.
شيكاغو (المؤلف–التاريخ)
Liu, Xiao, Haoyu Li, Jianghao Leng, Lin Wang, and Chao Sun. 2026. "GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving." https://omanscience.com/ar/articles/grava-grounded-reasoning-to-action-representation-and-learning-for-autonomous-driving.
هارفارد
Liu, X., Li, H., Leng, J., Wang, L. and Sun, C. (2026) 'GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving', Available at: https://omanscience.com/ar/articles/grava-grounded-reasoning-to-action-representation-and-learning-for-autonomous-driving.
فانكوفر
Liu X, Li H, Leng J, Wang L, Sun C. GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving. https://omanscience.com/ar/articles/grava-grounded-reasoning-to-action-representation-and-learning-for-autonomous-driving
IEEE
X. Liu, H. Li, J. Leng, L. Wang, and C. Sun, "GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving," https://omanscience.com/ar/articles/grava-grounded-reasoning-to-action-representation-and-learning-for-autonomous-driving.