Abstract

Post-training quantization (PTQ) has become a widely adopted technique for reducing the memory footprint and inference cost of large language models (LLMs). However, recent studies reveal that when applied to reasoning models, PTQ not only degrades reasoning performance but also exacerbates overthinking, leading to longer reasoning trajectories. These issues may offset the efficiency gains expected from lower-precision inference. Existing approaches mainly rely on complex optimization procedures. More recent lightweight inference strategies instead use predefined overthinking markers, limiting their adaptability across quantized models. To address these issues, we propose Reasoning Analysis and Token-level Inference Optimization (RATIO), a framework that identifies model-specific overthinking tokens and assigns each a tailored penalty. RATIO first introduces Quantization-aware Reasoning Behavior Analysis (QRBA) to identify overthinking tokens by analyzing discrepancies between full-precision and quantized models. It then adopts Token-Specific Penalty Determination (TSPD), which leverages full-precision guidance to derive token-specific penalties without additional training. Extensive experiments show that RATIO achieves a better accuracy-efficiency trade-off than existing token-level interventions. Specifically, RATIO achieves up to 9.8 points accuracy improvement and reduces chain-of-thought (CoT) length by up to 51.3% compared with quantized baselines. The code will be available at https://github.com/steven-bao1/RATIO.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Bao, C., Yan, X., Zhang, T., Chen, J., Zhang, S., & Zhang, Y. (2026). RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models. https://omanscience.com/en/articles/ratio-reasoning-analysis-and-token-level-inference-optimization-for-quantized-reasoning-models

MLA 9

Bao, Chengzhu, et al. "RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models." https://omanscience.com/en/articles/ratio-reasoning-analysis-and-token-level-inference-optimization-for-quantized-reasoning-models.

Chicago (author–date)

Bao, Chengzhu, Xianglong Yan, Tianao Zhang, Jiaqi Chen, Shaoqiu Zhang, and Yulun Zhang. 2026. "RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models." https://omanscience.com/en/articles/ratio-reasoning-analysis-and-token-level-inference-optimization-for-quantized-reasoning-models.

Harvard

Bao, C., Yan, X., Zhang, T., Chen, J., Zhang, S. and Zhang, Y. (2026) 'RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models', Available at: https://omanscience.com/en/articles/ratio-reasoning-analysis-and-token-level-inference-optimization-for-quantized-reasoning-models.

Vancouver

Bao C, Yan X, Zhang T, Chen J, Zhang S, Zhang Y. RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models. https://omanscience.com/en/articles/ratio-reasoning-analysis-and-token-level-inference-optimization-for-quantized-reasoning-models

IEEE

C. Bao, X. Yan, T. Zhang, J. Chen, S. Zhang, and Y. Zhang, "RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models," https://omanscience.com/en/articles/ratio-reasoning-analysis-and-token-level-inference-optimization-for-quantized-reasoning-models.