نسخة أولية وصول مفتوح
LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL
While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performin …