Abstract

Critic-free reinforcement fine-tuning (RFT) for agentic large language models is often done through GRPO-style methods, which compute a group baseline over repeated rollouts to reduce target variance. However, this setup is ill-suited to agents acting in stateful environments such as live services or security sandboxes, where repeated rollouts are impractical to obtain and aggressive updates entrench the noise of long, sparsely verified trajectories. We propose \textit{Follow the Winners} (FTW), a critic-free policy-learning algorithm that adapts the cross-entropy method to RFT, replacing group rollouts with an ordinal filter on replay-buffer samples that yields polynomial concentration in the order statistic of returns. We derive FTW through a control-as-inference lens, which also recovers GRPO and DPO as specific modelling choices, identifying GRPO as risk-neutral while DPO and FTW share a bounded risk-seeking offset that FTW controls. We identify this offset as an inherent trade-off of variance reduction through ordinal filters on samples, whereas a critic model induces a different trade-off between bias and variance. Scaled to agentic LLM post-training, FTW matches GRPO and PPO on Sokoban and Search-R1 baselines, showing a viable trade-off from a value model or group rollouts to CPU memory.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

de Vries, J. A., Lawrence, N. D., & Dai, Z. (2026). Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT. https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft

MLA 9

de Vries, Joery Ariën, et al. "Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT." https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft.

Chicago (author–date)

de Vries, Joery Ariën, Neil David Lawrence, and Zhenwen Dai. 2026. "Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT." https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft.

Harvard

de Vries, J. A., Lawrence, N. D. and Dai, Z. (2026) 'Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT', Available at: https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft.

Vancouver

de Vries JA, Lawrence ND, Dai Z. Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT. https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft

IEEE

J. A. de Vries, N. D. Lawrence, and Z. Dai, "Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT," https://omanscience.com/en/articles/follow-the-winners-conservative-policy-improvement-with-the-cross-entropy-method-for-critic-free-rft.