الملخص
Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wang, Y., Liu, Y., Tian, Q., Chen, X., Zhang, Z., Tu, Z., & Wang, R. (2026). Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving. https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving
MLA 9
Wang, Yiming, et al. "Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving." https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving.
شيكاغو (المؤلف–التاريخ)
Wang, Yiming, Yikang Liu, Qingyuan Tian, Xingyu Chen, Zhuosheng Zhang, Zhaopeng Tu, and Rui Wang. 2026. "Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving." https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving.
هارفارد
Wang, Y., Liu, Y., Tian, Q., Chen, X., Zhang, Z., Tu, Z. and Wang, R. (2026) 'Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving', Available at: https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving.
فانكوفر
Wang Y, Liu Y, Tian Q, Chen X, Zhang Z, Tu Z, et al. Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving. https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving
IEEE
Y. Wang, Y. Liu, Q. Tian, X. Chen, Z. Zhang, Z. Tu, and R. Wang, "Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving," https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving.