الملخص

Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges on reliable value estimation, a difficult task requiring the critic to both assess progress toward a correct solution and anticipate an evolving policy's future behavior; errors in either can compromise credit assignment and destabilize online training. In this paper, we revisit the standard state-only formulation of value estimation and propose $π$PPO, a self-privileged actor-critic framework. By reusing verified same-prompt rollouts as contrastive evidence, $π$PPO helps the critic assess intermediate reasoning against successful and failed attempts, while preserving standard policy optimization and the deployment interface. Experiments show that $π$PPO consistently improves value-estimation quality by a substantial margin and outperforms representative actor-critic and critic-free RLVR baselines on challenging mathematical reasoning benchmarks, while remaining effective even when paired with substantially smaller asymmetric critics.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Liang, K., Tang, C., Bai, C., Liu, W., Liu, Z., Zhang, Q., Yang, S., & Wu, Y. (2026). Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR. https://omanscience.com/ar/articles/privy-to-the-foil-recasting-value-estimation-with-a-self-privileged-critic-for-rlvr

MLA 9

Liang, Kun, et al. "Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR." https://omanscience.com/ar/articles/privy-to-the-foil-recasting-value-estimation-with-a-self-privileged-critic-for-rlvr.

شيكاغو (المؤلف–التاريخ)

Liang, Kun, Chenming Tang, Clive Bai, Weijie Liu, Zeyuan Liu, Qingyang Zhang, Saiyong Yang, and Yunfang Wu. 2026. "Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR." https://omanscience.com/ar/articles/privy-to-the-foil-recasting-value-estimation-with-a-self-privileged-critic-for-rlvr.

هارفارد

Liang, K., Tang, C., Bai, C., Liu, W., Liu, Z., Zhang, Q., Yang, S. and Wu, Y. (2026) 'Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR', Available at: https://omanscience.com/ar/articles/privy-to-the-foil-recasting-value-estimation-with-a-self-privileged-critic-for-rlvr.

فانكوفر

Liang K, Tang C, Bai C, Liu W, Liu Z, Zhang Q, et al. Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR. https://omanscience.com/ar/articles/privy-to-the-foil-recasting-value-estimation-with-a-self-privileged-critic-for-rlvr

IEEE

K. Liang, C. Tang, C. Bai, W. Liu, Z. Liu, Q. Zhang, S. Yang, and Y. Wu, "Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR," https://omanscience.com/ar/articles/privy-to-the-foil-recasting-value-estimation-with-a-self-privileged-critic-for-rlvr.