الباحثون

Navid Azizan

المنشورات 1

نسخة أولية وصول مفتوح

Score-Calibrated Flow for Sampling from Unnormalized Densities with Applications to Generative Online Reinforcement Learning

Zeyang Li, Yunan Wang, Risheek Garrepalli وآخرون · 2026

Diffusion and flow models provide expressive policy classes for online reinforcement learning (RL), enabling multimodal behaviors and improved performance. However, training these policies remains challenging: the critic specifies the desired policy as an unnormalized Boltzmann density but does not provide direct sampl …

المؤلفون المشاركون