الباحثون

Shuo Nie

المنشورات 1

نسخة أولية وصول مفتوح

GRPODropout: Less is More for Online Reinforcement Learning Rollouts

Hexuan Deng, Zihao Yan, Xuebo Liu وآخرون · 2026

Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such a …

المؤلفون المشاركون