الملخص
Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update. For compatible gradients, a cosine-dependent interpolation coordinates their contributions through a partially normalized reference while preserving the norm of their sum. We characterize this update as the unique solution of a spherical directional compromise. For conflicting gradients, projection follows the task's priorities. We evaluate the same compatible rule in helpfulness--safety alignment and correctness--cost optimization for mathematical reasoning. ORPG substantially improves average Useful and Harmless scores over the strongest external baseline on each axis. In mathematics, it achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods, with more accurate and shorter responses than the initial policy. Component comparisons and training dynamics show the larger contribution of compatible coordination and a complementary benefit from conflict handling. These results support gradient reconciliation for objectives with equal standing and for objectives with an explicit priority.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Fang, S., Zhao, Y., Tian, W., Lu, J., Zheng, Y., Wang, Y., & Qiu, X. (2026). ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients. https://omanscience.com/ar/articles/orpg-reconciling-multiple-reward-objectives-through-objective-wise-policy-gradients
MLA 9
Fang, Shicheng, et al. "ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients." https://omanscience.com/ar/articles/orpg-reconciling-multiple-reward-objectives-through-objective-wise-policy-gradients.
شيكاغو (المؤلف–التاريخ)
Fang, Shicheng, Yiwen Zhao, Wenbo Tian, Jiahao Lu, Yining Zheng, Yuxin Wang, and Xipeng Qiu. 2026. "ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients." https://omanscience.com/ar/articles/orpg-reconciling-multiple-reward-objectives-through-objective-wise-policy-gradients.
هارفارد
Fang, S., Zhao, Y., Tian, W., Lu, J., Zheng, Y., Wang, Y. and Qiu, X. (2026) 'ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients', Available at: https://omanscience.com/ar/articles/orpg-reconciling-multiple-reward-objectives-through-objective-wise-policy-gradients.
فانكوفر
Fang S, Zhao Y, Tian W, Lu J, Zheng Y, Wang Y, et al. ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients. https://omanscience.com/ar/articles/orpg-reconciling-multiple-reward-objectives-through-objective-wise-policy-gradients
IEEE
S. Fang, Y. Zhao, W. Tian, J. Lu, Y. Zheng, Y. Wang, and X. Qiu, "ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients," https://omanscience.com/ar/articles/orpg-reconciling-multiple-reward-objectives-through-objective-wise-policy-gradients.