Abstract
Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggregate gradient. We introduce SERA (Scale-Equalized Rollout Allocation), which redistributes a fixed rollout budget to approximately equalize these finite-rollout scaling factors. Building on our theoretical analysis of how finite rollouts distort prompt-wise likelihood gradients, we formulate the allocation as a fixed-budget max--min problem, derive a waterline solution to its continuous relaxation, and introduce a multiplicity correction to remove the additional prompt weighting induced by heterogeneous rollout counts. Experiments show stronger alignment with exact likelihood gradients in a controlled ImageNet setting and improved multi-sample solution coverage over MaxRL on maze navigation and mathematical reasoning under matched training rollout budgets.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Chen, Z., Xiong, F., Ren, H., Bai, X., Dai, Z., Chen, K., Zhang, Z., Wang, Z., & Cheng, Y. (2026). SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning. https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning
MLA 9
Chen, Zihao, et al. "SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning." https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning.
Chicago (author–date)
Chen, Zihao, Fanxiang Xiong, Hongran Ren, Xuefeng Bai, Zhongxiang Dai, Kehai Chen, Zhiguo Zhang, Zhiyong Wang, and Yu Cheng. 2026. "SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning." https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning.
Harvard
Chen, Z., Xiong, F., Ren, H., Bai, X., Dai, Z., Chen, K., Zhang, Z., Wang, Z. and Cheng, Y. (2026) 'SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning', Available at: https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning.
Vancouver
Chen Z, Xiong F, Ren H, Bai X, Dai Z, Chen K, et al. SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning. https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning
IEEE
Z. Chen, F. Xiong, H. Ren, X. Bai, Z. Dai, K. Chen, Z. Zhang, Z. Wang, and Y. Cheng, "SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning," https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning.