الملخص

Large language models (LLMs) are typically pretrained in high precision but increasingly deployed with low-precision post-training quantization (PTQ). Recent studies have shown that using weight averaging during pretraining can improve PTQ performance compared with learning-rate decay, suggesting that it might provide a simple way to improve the pretraining-to-quantization transition. But the mechanism behind weight averaging remains insufficiently explained. This leads to inconsistent and fragile performance gains, thereby preventing practitioners from applying such a technique confidently. As a response, we formulate weight averaging as a trade-off between retaining training progress and improving robustness under perturbation. We further derive a continuous family of averaging kernels that unifies conventional strategies and achieves the Pareto frontier between the two competing goals. Critically, a theoretical framework for performing weight averaging under PTQ is developed. It can be shown that coarser quantization is more susceptible to perturbations, whereas finer quantization could be less affected. Thus, our results could provide unified theoretical guidance for performing weight averaging under different PTQ conditions. Experiments validate both the predicted behavior and the proposed averaging strategy. Code is available at https://github.com/MOFA-LAB/weight-averaging-for-ptq.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Wang, H., Shen, T., Liu, Z., He, J., Zou, D., & Ma, Z. (2026). Understanding the Weight Averaging Mechanism in LLM Training for Post-Training Quantization. https://omanscience.com/ar/articles/understanding-the-weight-averaging-mechanism-in-llm-training-for-post-training-quantization

MLA 9

Wang, Hanzhang, et al. "Understanding the Weight Averaging Mechanism in LLM Training for Post-Training Quantization." https://omanscience.com/ar/articles/understanding-the-weight-averaging-mechanism-in-llm-training-for-post-training-quantization.

شيكاغو (المؤلف–التاريخ)

Wang, Hanzhang, Tianqi Shen, Zonglin Liu, Junze He, Difan Zou, and Ziye Ma. 2026. "Understanding the Weight Averaging Mechanism in LLM Training for Post-Training Quantization." https://omanscience.com/ar/articles/understanding-the-weight-averaging-mechanism-in-llm-training-for-post-training-quantization.

هارفارد

Wang, H., Shen, T., Liu, Z., He, J., Zou, D. and Ma, Z. (2026) 'Understanding the Weight Averaging Mechanism in LLM Training for Post-Training Quantization', Available at: https://omanscience.com/ar/articles/understanding-the-weight-averaging-mechanism-in-llm-training-for-post-training-quantization.

فانكوفر

Wang H, Shen T, Liu Z, He J, Zou D, Ma Z. Understanding the Weight Averaging Mechanism in LLM Training for Post-Training Quantization. https://omanscience.com/ar/articles/understanding-the-weight-averaging-mechanism-in-llm-training-for-post-training-quantization

IEEE

H. Wang, T. Shen, Z. Liu, J. He, D. Zou, and Z. Ma, "Understanding the Weight Averaging Mechanism in LLM Training for Post-Training Quantization," https://omanscience.com/ar/articles/understanding-the-weight-averaging-mechanism-in-llm-training-for-post-training-quantization.