Abstract
Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods are often unstable under meaning-preserving perturbations. Semantically equivalent paraphrases can induce substantial variability in predictive confidence, even for methods with formal guarantees, such as conformal prediction. To address this issue, we propose a paraphrase-aware UQ framework robust to semantic rewordings. Our approach trains a lightweight proxy model on LLM hidden states and aggregates its predictions across paraphrases to construct label-wise nonconformity scores. Under score exchangeability, conformal calibration retains marginal coverage. This guarantee can also hold under test-only rewording, provided that the paraphrase pipeline satisfies an additional distributional alignment condition. We evaluate three settings (normal, fully reworded, and semi-reworded) which apply rewording to neither dataset, both calibration and test datasets, or only the test dataset, respectively. Across seven multiple-choice QA benchmarks and multiple model families, our method produces compact prediction sets with empirical coverage generally near the nominal target, even in the semi-reworded setting. Ablation studies show that the learned proxy accounts for most of the reduction in set size, while paraphrase-augmented training and inference-time aggregation improve stability under rewording. Code is available at https://github.com/Raina-Xin/PA_Score.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Xin, J., Qiang, E., Zhu, Z., Li, X., Su, W. J., & Long, Q. (2026). Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification. https://omanscience.com/en/articles/conformal-prediction-with-paraphrase-aware-scoring-for-llm-uncertainty-quantification
MLA 9
Xin, Jiayi, et al. "Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification." https://omanscience.com/en/articles/conformal-prediction-with-paraphrase-aware-scoring-for-llm-uncertainty-quantification.
Chicago (author–date)
Xin, Jiayi, Evan Qiang, Zihan Zhu, Xiang Li, Weijie J. Su, and Qi Long. 2026. "Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification." https://omanscience.com/en/articles/conformal-prediction-with-paraphrase-aware-scoring-for-llm-uncertainty-quantification.
Harvard
Xin, J., Qiang, E., Zhu, Z., Li, X., Su, W. J. and Long, Q. (2026) 'Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification', Available at: https://omanscience.com/en/articles/conformal-prediction-with-paraphrase-aware-scoring-for-llm-uncertainty-quantification.
Vancouver
Xin J, Qiang E, Zhu Z, Li X, Su WJ, Long Q. Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification. https://omanscience.com/en/articles/conformal-prediction-with-paraphrase-aware-scoring-for-llm-uncertainty-quantification
IEEE
J. Xin, E. Qiang, Z. Zhu, X. Li, W. J. Su, and Q. Long, "Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification," https://omanscience.com/en/articles/conformal-prediction-with-paraphrase-aware-scoring-for-llm-uncertainty-quantification.