Preprint Open access
SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning
Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Un …