Abstract

Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such as reward modification and entropy/KL regularization, or through token-level reweighting. We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates. Under the same sampling budget, not all rollouts contribute positively to an update, and selectively excluding some can improve learning. To address this, we propose GRPODropout: before the standard update, we use a simple strategy that selectively removes a small number of high-probability positive-advantage rollouts and recenters the retained advantages. To motivate this design, we develop a rollout-level theoretical analysis that guides method design and threshold selection. The method changes only rollout usage, and adds negligible computational overhead. Experiments show higher accuracy than original GRPO and higher actor entropy while using fewer rollout samples for updates, illustrating "less is more." This work provides insight into RL rollout usage: removing some rollouts can improve performance. Code is available at https://github.com/hexuandeng/GRPODropout/.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Deng, H., Yan, Z., Liu, X., Nie, S., Wang, Y., Wang, C., Zhang, Z., Jiang, T., Xiao, Q., Zhang, J., & Zhang, M. (2026). GRPODropout: Less is More for Online Reinforcement Learning Rollouts. https://omanscience.com/en/articles/grpodropout-less-is-more-for-online-reinforcement-learning-rollouts

MLA 9

Deng, Hexuan, et al. "GRPODropout: Less is More for Online Reinforcement Learning Rollouts." https://omanscience.com/en/articles/grpodropout-less-is-more-for-online-reinforcement-learning-rollouts.

Chicago (author–date)

Deng, Hexuan, Zihao Yan, Xuebo Liu, Shuo Nie, Yue Wang, Chen Wang, Zhaohua Zhang, Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, and Min Zhang. 2026. "GRPODropout: Less is More for Online Reinforcement Learning Rollouts." https://omanscience.com/en/articles/grpodropout-less-is-more-for-online-reinforcement-learning-rollouts.

Harvard

Deng, H., Yan, Z., Liu, X., Nie, S., Wang, Y., Wang, C., Zhang, Z., Jiang, T., Xiao, Q., Zhang, J. and Zhang, M. (2026) 'GRPODropout: Less is More for Online Reinforcement Learning Rollouts', Available at: https://omanscience.com/en/articles/grpodropout-less-is-more-for-online-reinforcement-learning-rollouts.

Vancouver

Deng H, Yan Z, Liu X, Nie S, Wang Y, Wang C, et al. GRPODropout: Less is More for Online Reinforcement Learning Rollouts. https://omanscience.com/en/articles/grpodropout-less-is-more-for-online-reinforcement-learning-rollouts

IEEE

H. Deng, Z. Yan, X. Liu, S. Nie, Y. Wang, C. Wang, Z. Zhang, T. Jiang, Q. Xiao, J. Zhang, and M. Zhang, "GRPODropout: Less is More for Online Reinforcement Learning Rollouts," https://omanscience.com/en/articles/grpodropout-less-is-more-for-online-reinforcement-learning-rollouts.