الملخص
Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input correlations couple the errors of different groups, and useful code changes often involve many codes at once. We propose JARQ , a plug-in refinement that starts from any group-wise quantizer and alternates a joint least-squares fit of all group scales with bounded Babai proposals that move many codes of a group together on the current grid. The problem is a bilinear box-constrained mixed-integer least-squares problem; the solver is backpropagation-free, does not increase the layer-wise objective under exact scale solves, and keeps the host's bit width, groups, zero points, and inference cost. Across Llama-2, Llama-3, and Qwen models with RTN, GPTQ, OmniQuant, and AWQ hosts, JARQ lowers perplexity in 90 of 96 comparisons, cuts three-bit RTN perplexity by up to 36%, raises mean multiple-choice accuracy in 23 of 24 configurations, and improves QEP, QuaRot, and OJBKQ outputs, at under a minute per 7B block.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wang, X., Lyu, S., & Chang, X. W. (2026). JARQ: Joint Alternating Refinement for Quantization. https://omanscience.com/ar/articles/jarq-joint-alternating-refinement-for-quantization
MLA 9
Wang, Xinyu, et al. "JARQ: Joint Alternating Refinement for Quantization." https://omanscience.com/ar/articles/jarq-joint-alternating-refinement-for-quantization.
شيكاغو (المؤلف–التاريخ)
Wang, Xinyu, Sicheng Lyu, and Xiao-Wen Chang. 2026. "JARQ: Joint Alternating Refinement for Quantization." https://omanscience.com/ar/articles/jarq-joint-alternating-refinement-for-quantization.
هارفارد
Wang, X., Lyu, S. and Chang, X. W. (2026) 'JARQ: Joint Alternating Refinement for Quantization', Available at: https://omanscience.com/ar/articles/jarq-joint-alternating-refinement-for-quantization.
فانكوفر
Wang X, Lyu S, Chang XW. JARQ: Joint Alternating Refinement for Quantization. https://omanscience.com/ar/articles/jarq-joint-alternating-refinement-for-quantization
IEEE
X. Wang, S. Lyu, and X. W. Chang, "JARQ: Joint Alternating Refinement for Quantization," https://omanscience.com/ar/articles/jarq-joint-alternating-refinement-for-quantization.