الملخص
Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize knowledge acquired during vision-language pretraining for accurate and interpretable driving. However, VLAs can take only a limited number of frames as visual input due to the high token cost of an image, which is problematic for memory-dependent tasks such as determining the arrival order at all-way stops and long-horizon driving scene understanding. Existing solutions use latent vector memories accessed through cross-attention, which are neither interpretable nor portable. In this paper, we propose AD-Memo, a general-purpose VLA driving agent with language-based memory. The agent outputs memory as an extension of its Chain-of-Thought (CoT) to record surrounding objects critical to driving; this memory becomes part of the agent's future input. We curate memory-based datasets and train VLAs with a two-stage recipe: Supervised Fine-Tuning (SFT) and \textit{Da Capo}, a novel semi-closed-loop Reinforcement Learning (RL) algorithm which uses trajectory-level advantage for memory and step-level advantage for driving, leading to better credit assignment. Across scenarios such as all-way stops and general driving, AD-Memo improves driving quality, enables better question answering on driving scenes, and produces plug-and-play memory for other models.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Yan, K., Chen, X., Cao, Y., Naumann, A., Karkus, P., Wang, Y., Packer, J., Schwing, A., Wang, Y. X., Ivanovic, B., Luo, W., & Pavone, M. (2026). Vision-Language-Action Autonomous Driving Agent with Language-based Memory. https://omanscience.com/ar/articles/vision-language-action-autonomous-driving-agent-with-language-based-memory
MLA 9
Yan, Kai, et al. "Vision-Language-Action Autonomous Driving Agent with Language-based Memory." https://omanscience.com/ar/articles/vision-language-action-autonomous-driving-agent-with-language-based-memory.
شيكاغو (المؤلف–التاريخ)
Yan, Kai, Xiangyu Chen, Yulong Cao, Alex Naumann, Peter Karkus, Yan Wang, Jef Packer, Alex Schwing, Yu-Xiong Wang, Boris Ivanovic, Wenjie Luo, and Marco Pavone. 2026. "Vision-Language-Action Autonomous Driving Agent with Language-based Memory." https://omanscience.com/ar/articles/vision-language-action-autonomous-driving-agent-with-language-based-memory.
هارفارد
Yan, K., Chen, X., Cao, Y., Naumann, A., Karkus, P., Wang, Y., Packer, J., Schwing, A., Wang, Y. X., Ivanovic, B., Luo, W. and Pavone, M. (2026) 'Vision-Language-Action Autonomous Driving Agent with Language-based Memory', Available at: https://omanscience.com/ar/articles/vision-language-action-autonomous-driving-agent-with-language-based-memory.
فانكوفر
Yan K, Chen X, Cao Y, Naumann A, Karkus P, Wang Y, et al. Vision-Language-Action Autonomous Driving Agent with Language-based Memory. https://omanscience.com/ar/articles/vision-language-action-autonomous-driving-agent-with-language-based-memory
IEEE
K. Yan, X. Chen, Y. Cao, A. Naumann, P. Karkus, Y. Wang, J. Packer, A. Schwing, Y. X. Wang, B. Ivanovic, W. Luo, and M. Pavone, "Vision-Language-Action Autonomous Driving Agent with Language-based Memory," https://omanscience.com/ar/articles/vision-language-action-autonomous-driving-agent-with-language-based-memory.