Authors

Yiming Zong

Publications 2

Preprint Open access

CERO: Where and When to Allocate Rollouts for RL Post-Training

Yiming Zong, Yige Wang, Xing Hu et al. · 2026

Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and intro …

Co-authors