الملخص

As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but requires costly simulation for sensor-based policies. We propose OPTED (on-policy fine-tuning for end-to-end driving) which decouples reinforcement learning from the post-training of the end-to-end policy: a privileged teacher is trained using RL on vectorized inputs (HD-map and bounding boxes). This teacher then provides supervision to the pre-trained student during closed-loop post-training. We apply OPTED to two camera-based models, TransFuser and VaVAM, and fine-tune them in AlpaSim, using neural reconstructions (3DGS) of real driving logs. Driving scores increase by factors of 1.6$\times$ and 9.5$\times$, respectively. In controlled experiments OPTED matches closed-loop performance with approximately three orders of magnitude fewer simulator interactions than direct RL post-training, while staying closer to the human prior. Project page: https://01dami23.github.io/opted/

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Col, D. D., Igl, M., Karkus, P., Chitta, K., Ivanovic, B., Pavone, M., Schindler, K., & Sakaridis, C. (2026). OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher. https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher

MLA 9

Col, Damiano Da, et al. "OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher." https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher.

شيكاغو (المؤلف–التاريخ)

Col, Damiano Da, Maximilian Igl, Peter Karkus, Kashyap Chitta, Boris Ivanovic, Marco Pavone, Konrad Schindler, and Christos Sakaridis. 2026. "OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher." https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher.

هارفارد

Col, D. D., Igl, M., Karkus, P., Chitta, K., Ivanovic, B., Pavone, M., Schindler, K. and Sakaridis, C. (2026) 'OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher', Available at: https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher.

فانكوفر

Col DD, Igl M, Karkus P, Chitta K, Ivanovic B, Pavone M, et al. OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher. https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher

IEEE

D. D. Col, M. Igl, P. Karkus, K. Chitta, B. Ivanovic, M. Pavone, K. Schindler, and C. Sakaridis, "OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher," https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher.