Abstract

Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggregate gradient. We introduce SERA (Scale-Equalized Rollout Allocation), which redistributes a fixed rollout budget to approximately equalize these finite-rollout scaling factors. Building on our theoretical analysis of how finite rollouts distort prompt-wise likelihood gradients, we formulate the allocation as a fixed-budget max--min problem, derive a waterline solution to its continuous relaxation, and introduce a multiplicity correction to remove the additional prompt weighting induced by heterogeneous rollout counts. Experiments show stronger alignment with exact likelihood gradients in a controlled ImageNet setting and improved multi-sample solution coverage over MaxRL on maze navigation and mathematical reasoning under matched training rollout budgets.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Chen, Z., Xiong, F., Ren, H., Bai, X., Dai, Z., Chen, K., Zhang, Z., Wang, Z., & Cheng, Y. (2026). SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning. https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning

MLA 9

Chen, Zihao, et al. "SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning." https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning.

Chicago (author–date)

Chen, Zihao, Fanxiang Xiong, Hongran Ren, Xuefeng Bai, Zhongxiang Dai, Kehai Chen, Zhiguo Zhang, Zhiyong Wang, and Yu Cheng. 2026. "SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning." https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning.

Harvard

Chen, Z., Xiong, F., Ren, H., Bai, X., Dai, Z., Chen, K., Zhang, Z., Wang, Z. and Cheng, Y. (2026) 'SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning', Available at: https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning.

Vancouver

Chen Z, Xiong F, Ren H, Bai X, Dai Z, Chen K, et al. SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning. https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning

IEEE

Z. Chen, F. Xiong, H. Ren, X. Bai, Z. Dai, K. Chen, Z. Zhang, Z. Wang, and Y. Cheng, "SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning," https://omanscience.com/en/articles/sera-scale-equalized-rollout-allocation-for-maximum-likelihood-reinforcement-learning.