Abstract
Quantization-aware training (QAT) leverages lower-precision arithmetic to reduce the cost of LLM deployment, but aggressive quantization degrades final model performance. A common remedy is mixed-precision training, in which high precision is assigned to some of the layers to maintain performance while keeping the cost constrained. This approach then requires precision assignments for model layers during training. We provide a new approach, called Q-PACE, consisting of a second-order sensitivity model that predicts the loss increase as a sum of quantization noise MSE weighted by per-layer curvature coefficients. During training, we periodically re-compute these coefficients using perturbations across layers, and re-assign precision. Pretraining and supervised fine-tuning experiments on LLMs of up to 4B parameters show that Q-PACE consistently improves over existing mixed-precision training recipes, and achieves comparable loss at substantially lower total memory budgets. We further find that quantization sensitivity is highly predictable by depth and layer type, and its stability during training allows for infrequent, cheap recalibration.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Volkova, A., Ansaripour, M., Schultheis, E., Lampert, C. H., & Alistarh, D. (2026). Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training. https://omanscience.com/en/articles/q-pace-dynamic-precision-allocation-for-quantization-aware-training
MLA 9
Volkova, Alexandra, et al. "Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training." https://omanscience.com/en/articles/q-pace-dynamic-precision-allocation-for-quantization-aware-training.
Chicago (author–date)
Volkova, Alexandra, Matin Ansaripour, Erik Schultheis, Christoph H. Lampert, and Dan Alistarh. 2026. "Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training." https://omanscience.com/en/articles/q-pace-dynamic-precision-allocation-for-quantization-aware-training.
Harvard
Volkova, A., Ansaripour, M., Schultheis, E., Lampert, C. H. and Alistarh, D. (2026) 'Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training', Available at: https://omanscience.com/en/articles/q-pace-dynamic-precision-allocation-for-quantization-aware-training.
Vancouver
Volkova A, Ansaripour M, Schultheis E, Lampert CH, Alistarh D. Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training. https://omanscience.com/en/articles/q-pace-dynamic-precision-allocation-for-quantization-aware-training
IEEE
A. Volkova, M. Ansaripour, E. Schultheis, C. H. Lampert, and D. Alistarh, "Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training," https://omanscience.com/en/articles/q-pace-dynamic-precision-allocation-for-quantization-aware-training.