الملخص
Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is calibrated. The resulting models are systematically overconfident. Recent methods train calibration inside the RLVR loop by having the model state a numerical confidence alongside its answer, but they all obtain the confidence by sampling it as text. This choice imposes two costs: a sampled confidence introduces variance and in practice collapses to a handful of distinct values, and sampling makes the confidence non-differentiable, forcing the calibration loss through a scalar reward. We propose CREDO (Confidence REaDOut) to replace sampling with a deterministic readout. While RLVR optimizes correctness, CREDO reads the confidence from a dedicated token pair in the model's output distribution and trains it by differentiable regression. CREDO further turns the trained confidence into a signal for accuracy, weighting rollouts by how far confidence and outcome disagree, so that accuracy and calibration improve together. Across mathematical and code reasoning, CREDO attains the best accuracy and calibration, and the gains extend to abstention and selective prediction.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Fan, C., Gao, C., Zhang, G., Shen, L., Gong, Y., Wang, J., Wang, J., Wang, D., Liu, Y., Feng, F., & He, X. (2026). Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts. https://omanscience.com/ar/articles/beyond-verbalized-confidence-calibrating-reasoners-with-differentiable-readouts
MLA 9
Fan, Chenxiao, et al. "Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts." https://omanscience.com/ar/articles/beyond-verbalized-confidence-calibrating-reasoners-with-differentiable-readouts.
شيكاغو (المؤلف–التاريخ)
Fan, Chenxiao, Chongming Gao, Gangyi Zhang, Leyang Shen, Yaxin Gong, Jiamin Wang, Jiakai Wang, Dong Wang, Yang Liu, Fuli Feng, and Xiangnan He. 2026. "Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts." https://omanscience.com/ar/articles/beyond-verbalized-confidence-calibrating-reasoners-with-differentiable-readouts.
هارفارد
Fan, C., Gao, C., Zhang, G., Shen, L., Gong, Y., Wang, J., Wang, J., Wang, D., Liu, Y., Feng, F. and He, X. (2026) 'Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts', Available at: https://omanscience.com/ar/articles/beyond-verbalized-confidence-calibrating-reasoners-with-differentiable-readouts.
فانكوفر
Fan C, Gao C, Zhang G, Shen L, Gong Y, Wang J, et al. Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts. https://omanscience.com/ar/articles/beyond-verbalized-confidence-calibrating-reasoners-with-differentiable-readouts
IEEE
C. Fan, C. Gao, G. Zhang, L. Shen, Y. Gong, J. Wang, J. Wang, D. Wang, Y. Liu, F. Feng, and X. He, "Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts," https://omanscience.com/ar/articles/beyond-verbalized-confidence-calibrating-reasoners-with-differentiable-readouts.