Abstract
Critic-free reinforcement fine-tuning (RFT) for agentic large language models is often done through GRPO-style methods, which compute a group baseline over repeated rollouts to reduce target variance. However, this setup is ill-suited to agents acting in stateful environments such as live services or security sandboxes, where repeated rollouts are impractical to obtain and aggressive updates entrench the noise of long, sparsely verified trajectories. We propose \textit{Follow the Winners} (FTW), a critic-free policy-learning algorithm that adapts the cross-entropy method to RFT, replacing group rollouts with an ordinal filter on replay-buffer samples that yields polynomial concentration in the order statistic of returns. We derive FTW through a control-as-inference lens, which also recovers GRPO and DPO as specific modelling choices, identifying GRPO as risk-neutral while DPO and FTW share a bounded risk-seeking offset that FTW controls. We identify this offset as an inherent trade-off of variance reduction through ordinal filters on samples, whereas a critic model induces a different trade-off between bias and variance. Scaled to agentic LLM post-training, FTW matches GRPO and PPO on Sokoban and Search-R1 baselines, showing a viable trade-off from a value model or group rollouts to CPU memory.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
de Vries, J. A., Lawrence, N. D., & Dai, Z. (2026). Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT. https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft
MLA 9
de Vries, Joery Ariën, et al. "Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT." https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft.
Chicago (author–date)
de Vries, Joery Ariën, Neil David Lawrence, and Zhenwen Dai. 2026. "Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT." https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft.
Harvard
de Vries, J. A., Lawrence, N. D. and Dai, Z. (2026) 'Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT', Available at: https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft.
Vancouver
de Vries JA, Lawrence ND, Dai Z. Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT. https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft
IEEE
J. A. de Vries, N. D. Lawrence, and Z. Dai, "Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT," https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft.