Preprint Open access
Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models
Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer, not how likely the …