نسخة أولية وصول مفتوح
GRPODropout: Less is More for Online Reinforcement Learning Rollouts
Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such a …