Abstract

Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pruning strategy to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. 2) Then, we design an adaptive rollout sampling mechanism to dynamically adjust the sampling scale across different training stages based on historical pruning distributions, balancing exploration adequacy and computational efficiency. Experiments demonstrate that FastRL can be seamlessly integrated into GRPO, DAPO, and GSPO variants, achieving an average 2.07$\times$ training speedup on Geometry3K and GeoQA8K-R1V, along with an approximately 1.64\% improvement in average accuracy on visual reasoning benchmarks. Source codes will be available at https://github.com/Nicozwy/FastRL.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Yang, J., Yang, Z., Zhang, X., Chen, D., Chen, X., Su, T., Lu, H., Guan, Q., Tang, K., & Wang, C. (2026). Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO. https://omanscience.com/en/articles/learn-from-the-gap-differential-aware-advantage-pruning-with-adaptive-rollout-sampling-for-grpo

MLA 9

Yang, Jiahua, et al. "Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO." https://omanscience.com/en/articles/learn-from-the-gap-differential-aware-advantage-pruning-with-adaptive-rollout-sampling-for-grpo.

Chicago (author–date)

Yang, Jiahua, Zhiwei Yang, Xianpeng Zhang, Dongyu Chen, Xing Chen, Tianhuang Su, Haonan Lu, Quanlong Guan, Kai Tang, and Chuangchuang Wang. 2026. "Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO." https://omanscience.com/en/articles/learn-from-the-gap-differential-aware-advantage-pruning-with-adaptive-rollout-sampling-for-grpo.

Harvard

Yang, J., Yang, Z., Zhang, X., Chen, D., Chen, X., Su, T., Lu, H., Guan, Q., Tang, K. and Wang, C. (2026) 'Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO', Available at: https://omanscience.com/en/articles/learn-from-the-gap-differential-aware-advantage-pruning-with-adaptive-rollout-sampling-for-grpo.

Vancouver

Yang J, Yang Z, Zhang X, Chen D, Chen X, Su T, et al. Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO. https://omanscience.com/en/articles/learn-from-the-gap-differential-aware-advantage-pruning-with-adaptive-rollout-sampling-for-grpo

IEEE

J. Yang, Z. Yang, X. Zhang, D. Chen, X. Chen, T. Su, H. Lu, Q. Guan, K. Tang, and C. Wang, "Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO," https://omanscience.com/en/articles/learn-from-the-gap-differential-aware-advantage-pruning-with-adaptive-rollout-sampling-for-grpo.