نسخة أولية وصول مفتوح
Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving
Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distribution shift: policy optimization may favor actions that are weakly supported by the offl …