[
    {
        "id": "osp-22144",
        "type": "article-journal",
        "title": "ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients",
        "author": [
            {
                "family": "Fang",
                "given": "Shicheng"
            },
            {
                "family": "Zhao",
                "given": "Yiwen"
            },
            {
                "family": "Tian",
                "given": "Wenbo"
            },
            {
                "family": "Lu",
                "given": "Jiahao"
            },
            {
                "family": "Zheng",
                "given": "Yining"
            },
            {
                "family": "Wang",
                "given": "Yuxin"
            },
            {
                "family": "Qiu",
                "given": "Xipeng"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/orpg-reconciling-multiple-reward-objectives-through-objective-wise-policy-gradients",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update. For compatible gradients, a cosine-dependent interpolation coordinates their contributions through a partially normalized reference while preserving the norm of their sum. We characterize this update as the unique solution of a spherical directional compromise. For conflicting gradients, projection follows the task's priorities. We evaluate the same compatible rule in helpfulness--safety alignment and correctness--cost optimization for mathematical reasoning. ORPG substantially improves average Useful and Harmless scores over the strongest external baseline on each axis. In mathematics, it achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods, with more accurate and shorter responses than the initial policy. Component comparisons and training dynamics show the larger contribution of compatible coordination and a complementary benefit from conflict handling. These results support gradient reconciliation for objectives with equal standing and for objectives with an explicit priority."
    }
]