الملخص

Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Wang, Y., Liu, Y., Tian, Q., Chen, X., Zhang, Z., Tu, Z., & Wang, R. (2026). Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving. https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving

MLA 9

Wang, Yiming, et al. "Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving." https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving.

شيكاغو (المؤلف–التاريخ)

Wang, Yiming, Yikang Liu, Qingyuan Tian, Xingyu Chen, Zhuosheng Zhang, Zhaopeng Tu, and Rui Wang. 2026. "Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving." https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving.

هارفارد

Wang, Y., Liu, Y., Tian, Q., Chen, X., Zhang, Z., Tu, Z. and Wang, R. (2026) 'Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving', Available at: https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving.

فانكوفر

Wang Y, Liu Y, Tian Q, Chen X, Zhang Z, Tu Z, et al. Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving. https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving

IEEE

Y. Wang, Y. Liu, Q. Tian, X. Chen, Z. Zhang, Z. Tu, and R. Wang, "Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving," https://omanscience.com/ar/articles/group-marginalized-self-rewarding-rl-drives-zero-label-self-evolving.