Abstract
Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself. Quantization errors can therefore move the model into states that are absent from offline recovery data. We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses. At each visited pre- fix, a frozen full-precision teacher provides a sampled reverse-KL training signal. On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively. The results suggest that student-visited states provide a useful recovery signal beyond fixed-completion training, particularly at three bits.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Wang, W., Li, H., Gu, Y., & Yang, H. (2026). OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models. https://omanscience.com/en/articles/onlineqat-on-policy-distillation-for-ultra-low-bit-large-language-models
MLA 9
Wang, Wenjun, et al. "OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models." https://omanscience.com/en/articles/onlineqat-on-policy-distillation-for-ultra-low-bit-large-language-models.
Chicago (author–date)
Wang, Wenjun, Heng Li, Yanggan Gu, and Hongxia Yang. 2026. "OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models." https://omanscience.com/en/articles/onlineqat-on-policy-distillation-for-ultra-low-bit-large-language-models.
Harvard
Wang, W., Li, H., Gu, Y. and Yang, H. (2026) 'OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models', Available at: https://omanscience.com/en/articles/onlineqat-on-policy-distillation-for-ultra-low-bit-large-language-models.
Vancouver
Wang W, Li H, Gu Y, Yang H. OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models. https://omanscience.com/en/articles/onlineqat-on-policy-distillation-for-ultra-low-bit-large-language-models
IEEE
W. Wang, H. Li, Y. Gu, and H. Yang, "OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models," https://omanscience.com/en/articles/onlineqat-on-policy-distillation-for-ultra-low-bit-large-language-models.