الملخص
As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but requires costly simulation for sensor-based policies. We propose OPTED (on-policy fine-tuning for end-to-end driving) which decouples reinforcement learning from the post-training of the end-to-end policy: a privileged teacher is trained using RL on vectorized inputs (HD-map and bounding boxes). This teacher then provides supervision to the pre-trained student during closed-loop post-training. We apply OPTED to two camera-based models, TransFuser and VaVAM, and fine-tune them in AlpaSim, using neural reconstructions (3DGS) of real driving logs. Driving scores increase by factors of 1.6$\times$ and 9.5$\times$, respectively. In controlled experiments OPTED matches closed-loop performance with approximately three orders of magnitude fewer simulator interactions than direct RL post-training, while staying closer to the human prior. Project page: https://01dami23.github.io/opted/
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Col, D. D., Igl, M., Karkus, P., Chitta, K., Ivanovic, B., Pavone, M., Schindler, K., & Sakaridis, C. (2026). OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher. https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher
MLA 9
Col, Damiano Da, et al. "OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher." https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher.
شيكاغو (المؤلف–التاريخ)
Col, Damiano Da, Maximilian Igl, Peter Karkus, Kashyap Chitta, Boris Ivanovic, Marco Pavone, Konrad Schindler, and Christos Sakaridis. 2026. "OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher." https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher.
هارفارد
Col, D. D., Igl, M., Karkus, P., Chitta, K., Ivanovic, B., Pavone, M., Schindler, K. and Sakaridis, C. (2026) 'OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher', Available at: https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher.
فانكوفر
Col DD, Igl M, Karkus P, Chitta K, Ivanovic B, Pavone M, et al. OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher. https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher
IEEE
D. D. Col, M. Igl, P. Karkus, K. Chitta, B. Ivanovic, M. Pavone, K. Schindler, and C. Sakaridis, "OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher," https://omanscience.com/ar/articles/opted-on-policy-fine-tuning-for-end-to-end-driving-using-a-render-free-teacher.