Abstract

Recent VLM-based autonomous driving planners adopt GRPO-style reinforcement learning to optimize driving performance. However, existing GRPO recipes either optimize driving efficiency, risking progress-seeking but unsafe behavior, or enforce early safety constraints, leading to overly conservative behavior; both require lengthy training. To solve these problems, we first reveal two distinct RL regimes: a progress regime (Run-GRPO) that aggressively explores high progress, and a safety regime (Walk-GRPO) that restores safety under stable progress. Based on this finding, we propose $\textit{Run-then-Walk}$, a simple yet effective two-stage reward scheduling strategy for GRPO, achieving both better performance and faster convergence. Unlike one-stage RL, which may focus on progress, safety, or a mixture of both within a single training phase, this schedule explicitly separates progress discovery from safety repair. In the $\textit{Run}$ phase, we focus on progress, allowing the policy to escape the conservative bias and discover high-progress modes. In the subsequent $\textit{Walk}$ phase, we introduce endpoint and safety strategy to repair unsafe behaviors from the Run phase. This reversed schedule overcomes the conservatism of Walk-first methods and the unsafe progress-seeking of joint optimization. We validate it with various VLM-based planners on multiple benchmarks: NAVSIMv1, NAVSIMv2, Navhard, and nuScenes. Extensive experiments demonstrate improved driving performance while requiring 40--50\% fewer RL training epochs than the baselines. Code is available at https://github.com/haha-yuki-haha/AutoDrive-P3_with_Run-then-walk.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Ye, Y., Sun, S., Lin, J., Zhao, J., Peng, C., Zheng, W., Liu, G., Zhao, T., & Gao, W. (2026). Sometimes You Gotta Run Before You Can Walk: Run-then-Walk Scheduling Strategy for VLM Autonomous Driving. https://omanscience.com/en/articles/sometimes-you-gotta-run-before-you-can-walk-run-then-walk-scheduling-strategy-for-vlm-autonomous-driving

MLA 9

Ye, Yuqi, et al. "Sometimes You Gotta Run Before You Can Walk: Run-then-Walk Scheduling Strategy for VLM Autonomous Driving." https://omanscience.com/en/articles/sometimes-you-gotta-run-before-you-can-walk-run-then-walk-scheduling-strategy-for-vlm-autonomous-driving.

Chicago (author–date)

Ye, Yuqi, Shangkun Sun, Junhong Lin, Jiayi Zhao, Changhao Peng, Wei Zheng, Guoqing Liu, Tiesong Zhao, and Wei Gao. 2026. "Sometimes You Gotta Run Before You Can Walk: Run-then-Walk Scheduling Strategy for VLM Autonomous Driving." https://omanscience.com/en/articles/sometimes-you-gotta-run-before-you-can-walk-run-then-walk-scheduling-strategy-for-vlm-autonomous-driving.

Harvard

Ye, Y., Sun, S., Lin, J., Zhao, J., Peng, C., Zheng, W., Liu, G., Zhao, T. and Gao, W. (2026) 'Sometimes You Gotta Run Before You Can Walk: Run-then-Walk Scheduling Strategy for VLM Autonomous Driving', Available at: https://omanscience.com/en/articles/sometimes-you-gotta-run-before-you-can-walk-run-then-walk-scheduling-strategy-for-vlm-autonomous-driving.

Vancouver

Ye Y, Sun S, Lin J, Zhao J, Peng C, Zheng W, et al. Sometimes You Gotta Run Before You Can Walk: Run-then-Walk Scheduling Strategy for VLM Autonomous Driving. https://omanscience.com/en/articles/sometimes-you-gotta-run-before-you-can-walk-run-then-walk-scheduling-strategy-for-vlm-autonomous-driving

IEEE

Y. Ye, S. Sun, J. Lin, J. Zhao, C. Peng, W. Zheng, G. Liu, T. Zhao, and W. Gao, "Sometimes You Gotta Run Before You Can Walk: Run-then-Walk Scheduling Strategy for VLM Autonomous Driving," https://omanscience.com/en/articles/sometimes-you-gotta-run-before-you-can-walk-run-then-walk-scheduling-strategy-for-vlm-autonomous-driving.