الباحثون

Yonghoon Dong

المنشورات 1

نسخة أولية وصول مفتوح

Q-Learning with Scalar Adjoint Matching

Yonghoon Dong, Minsung Yoon, Jaehyuk Kim وآخرون · 2026

Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint …

المؤلفون المشاركون