Preprint Open access
OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models
Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by its …