الملخص

Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning. First, we find that instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact: keeping policy updates small and low in variance restores stable convergence. Second, a well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. Its predictions provide outcome-derived, dense, per-prefix learning signals that, during policy optimization, require neither completed rollouts, step-level annotations, nor external reward labels. Building on this insight, we introduce Reward-Free Policy Optimization (RFPO), which repurposes a single calibrated, frozen critic as a rollout-level reward, a value baseline for generalized advantage estimation, and a success forecaster for unfinished prefixes. We further show that binarizing the debiased score stops the policy from exploiting the critic's length bias. Binarized, RFPO matches supervised PPO without a single label in the training loop, while cutting compute and memory overhead. This makes RFPO well suited to long-horizon reasoning tasks, where outcomes arrive late and generation dominates cost: because rollouts can be rewarded before they finish, training no longer has to pay for waiting on every trajectory to complete. Our findings challenge the prevailing critic-free paradigm and establish critic-based, reward-free optimization as a scalable and computationally efficient path for LLM post-training.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Li, H., Li, X., Wu, C., Mammar, S., Danoy, G., & Bouvry, P. (2026). Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training. https://omanscience.com/ar/articles/unlocking-the-critic-reward-free-policy-optimization-for-llm-post-training

MLA 9

Li, Hongyang, et al. "Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training." https://omanscience.com/ar/articles/unlocking-the-critic-reward-free-policy-optimization-for-llm-post-training.

شيكاغو (المؤلف–التاريخ)

Li, Hongyang, Xiao Li, Caesar Wu, Said Mammar, Grégoire Danoy, and Pascal Bouvry. 2026. "Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training." https://omanscience.com/ar/articles/unlocking-the-critic-reward-free-policy-optimization-for-llm-post-training.

هارفارد

Li, H., Li, X., Wu, C., Mammar, S., Danoy, G. and Bouvry, P. (2026) 'Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training', Available at: https://omanscience.com/ar/articles/unlocking-the-critic-reward-free-policy-optimization-for-llm-post-training.

فانكوفر

Li H, Li X, Wu C, Mammar S, Danoy G, Bouvry P. Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training. https://omanscience.com/ar/articles/unlocking-the-critic-reward-free-policy-optimization-for-llm-post-training

IEEE

H. Li, X. Li, C. Wu, S. Mammar, G. Danoy, and P. Bouvry, "Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training," https://omanscience.com/ar/articles/unlocking-the-critic-reward-free-policy-optimization-for-llm-post-training.