الملخص

4D radar provides geometric and motion cues that complement visual semantics, but integrating it into vision-language-action (VLA) models requires both radar--language alignment for semantic reasoning and explicit use of radar measurements for trajectory refinement and selection. To support these capabilities, we construct Cap4DR with 86,016 radar-image-text samples for alignment pretraining and OmniHD-QA with 520,161 question-answer pairs for instruction tuning across scene description, key-object reasoning, occupancy understanding, and trajectory planning. Building on these datasets, we propose RCVLA, a radar-camera VLA framework consisting of a radar-grounded semantic reasoning stage (RCVLA-Sem) and a trajectory arbitration stage (RCVLA-Phys). RCVLA-Sem performs gated bidirectional interaction between camera and radar tokens for driving question answering and reference trajectory generation, while auxiliary heads provide object and occupancy queries. RCVLA-Phys refines reference-guided trajectory candidates through truncated diffusion conditioned on these queries and cluster-level radar measurements, then calibrates candidate scores using radar-derived time-to-collision risk. On OmniHD-QA, RCVLA-Sem improves CIDEr by 9.92 points and reduces key-object velocity error by $21.9\%$ relative to OmniDrive. RCVLA-Phys further reduces average L2 error from $0.348$ to $0.259\,\mathrm{m}$ and average open-loop collision rate from $0.576\%$ to $0.175\%$ relative to RCVLA-Sem. Ablation studies further show that language-aligned radar tokens improve semantic reasoning, while cluster-level radar measurements and risk calibration improve trajectory arbitration. Code will be released.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Zheng, L., Bai, X., Luo, Y., Guan, R., Liu, M., Wei, Z., Shen, H. L., Zhu, X., & Ma, Z. (2026). RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving. https://omanscience.com/ar/articles/rcvla-4d-radar-grounded-semantic-reasoning-and-trajectory-arbitration-for-autonomous-driving

MLA 9

Zheng, Lianqing, et al. "RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving." https://omanscience.com/ar/articles/rcvla-4d-radar-grounded-semantic-reasoning-and-trajectory-arbitration-for-autonomous-driving.

شيكاغو (المؤلف–التاريخ)

Zheng, Lianqing, Xiaokai Bai, Yixuan Luo, Runwei Guan, Minghao Liu, Zhiqiang Wei, Hui-liang Shen, Xichan Zhu, and Zhixiong Ma. 2026. "RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving." https://omanscience.com/ar/articles/rcvla-4d-radar-grounded-semantic-reasoning-and-trajectory-arbitration-for-autonomous-driving.

هارفارد

Zheng, L., Bai, X., Luo, Y., Guan, R., Liu, M., Wei, Z., Shen, H. L., Zhu, X. and Ma, Z. (2026) 'RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving', Available at: https://omanscience.com/ar/articles/rcvla-4d-radar-grounded-semantic-reasoning-and-trajectory-arbitration-for-autonomous-driving.

فانكوفر

Zheng L, Bai X, Luo Y, Guan R, Liu M, Wei Z, et al. RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving. https://omanscience.com/ar/articles/rcvla-4d-radar-grounded-semantic-reasoning-and-trajectory-arbitration-for-autonomous-driving

IEEE

L. Zheng, X. Bai, Y. Luo, R. Guan, M. Liu, Z. Wei, H. L. Shen, X. Zhu, and Z. Ma, "RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving," https://omanscience.com/ar/articles/rcvla-4d-radar-grounded-semantic-reasoning-and-trajectory-arbitration-for-autonomous-driving.