Preprint Open access
Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving
Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group cont …