الباحثون

Yiwen Zhao

المنشورات 1

نسخة أولية وصول مفتوح

ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients

Shicheng Fang, Yiwen Zhao, Wenbo Tian وآخرون · 2026

Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients int …

المؤلفون المشاركون