Abstract
Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident than their corresponding pretrained language models (PLMs). Because task-specific labeled calibration data can be costly or unavailable, the corresponding PLM provides a natural label-free reference for post-hoc calibration. Prior agreement-gated PLM-referenced calibration fits a scalar temperature using only examples on which the PoLM and its PLM reference agree, excluding disagreement examples because direct alignment can drive the fitted temperature excessively high and induce under-confidence. We revisit this binary treatment. A controlled reintroduction diagnostic reveals a non monotonic aggregate effect: admitting a moderate fraction of disagreement examples can improve calibration, whereas the benefit diminishes as unit weight inclusion approaches the full disagreement set. We introduce SupportCal, a label-free post-hoc method that retains agreement examples at unit weight and assigns disagreement examples continuous weights based on the own-base PLM's relative support and corroboration from pretrained references selected from a size-compatible candidate pool. We further characterize when the resulting weighted objective admits a finite optimal temperature. Across MedMCQA and MathQA, SupportCal yields lower mean ECE than the agreement-only baseline for nearly all evaluated target-model configurations; supplementary TweetEval Sentiment results show the same pattern on a fixed-label classification task.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Luo, L., Lin, L., Shi, D., Chen, F., Hernández-Lobato, J. M., & Gao, J. (2026). SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration. https://omanscience.com/en/articles/supportcal-label-free-calibration-of-post-trained-llms-via-reference-support-and-corroboration
MLA 9
Luo, Linhan, et al. "SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration." https://omanscience.com/en/articles/supportcal-label-free-calibration-of-post-trained-llms-via-reference-support-and-corroboration.
Chicago (author–date)
Luo, Linhan, Lequan Lin, Dai Shi, Feng Chen, José Miguel Hernández-Lobato, and Junbin Gao. 2026. "SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration." https://omanscience.com/en/articles/supportcal-label-free-calibration-of-post-trained-llms-via-reference-support-and-corroboration.
Harvard
Luo, L., Lin, L., Shi, D., Chen, F., Hernández-Lobato, J. M. and Gao, J. (2026) 'SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration', Available at: https://omanscience.com/en/articles/supportcal-label-free-calibration-of-post-trained-llms-via-reference-support-and-corroboration.
Vancouver
Luo L, Lin L, Shi D, Chen F, Hernández-Lobato JM, Gao J. SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration. https://omanscience.com/en/articles/supportcal-label-free-calibration-of-post-trained-llms-via-reference-support-and-corroboration
IEEE
L. Luo, L. Lin, D. Shi, F. Chen, J. M. Hernández-Lobato, and J. Gao, "SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration," https://omanscience.com/en/articles/supportcal-label-free-calibration-of-post-trained-llms-via-reference-support-and-corroboration.