Abstract

Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds. This paper presents G$^2$PTQ, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective. By refreshing gradient and Hessian estimates before quantizing each Transformer block, G$^2$PTQ avoids the staleness of prior global methods. Furthermore, to stabilize the exact first-order compensation, we introduce a trust-region scaling mechanism that dynamically bounds the gradient step to prevent exploding weight updates. Finally, we derive efficient implementations for block-wise Hessian approximation and exact gradient compensation. Experimental results on various model families and bit-widths demonstrate that G$^2$PTQ enables better alignment with the full-precision model, outperforming state-of-the-art baselines. Code is available at: https://github.com/G2PTQ/G2PTQ.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Liu, R., Bai, H., Sun, Y., Zhang, Q., Cai, W., Hao, Y., Wang, F., Zhong, W., Wang, Z., Yang, T., & Zhou, X. (2026). G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation. https://omanscience.com/en/articles/g-2-ptq-improving-llm-post-training-quantization-with-generalized-gradient-compensation

MLA 9

Liu, Ruikang, et al. "G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation." https://omanscience.com/en/articles/g-2-ptq-improving-llm-post-training-quantization-with-generalized-gradient-compensation.

Chicago (author–date)

Liu, Ruikang, Haoli Bai, Yuxuan Sun, Qian Zhang, Wenzheng Cai, Yanqi Hao, Feiyu Wang, Weidong Zhong, Zhuang Wang, Tong Yang, and Xiangsheng Zhou. 2026. "G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation." https://omanscience.com/en/articles/g-2-ptq-improving-llm-post-training-quantization-with-generalized-gradient-compensation.

Harvard

Liu, R., Bai, H., Sun, Y., Zhang, Q., Cai, W., Hao, Y., Wang, F., Zhong, W., Wang, Z., Yang, T. and Zhou, X. (2026) 'G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation', Available at: https://omanscience.com/en/articles/g-2-ptq-improving-llm-post-training-quantization-with-generalized-gradient-compensation.

Vancouver

Liu R, Bai H, Sun Y, Zhang Q, Cai W, Hao Y, et al. G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation. https://omanscience.com/en/articles/g-2-ptq-improving-llm-post-training-quantization-with-generalized-gradient-compensation

IEEE

R. Liu, H. Bai, Y. Sun, Q. Zhang, W. Cai, Y. Hao, F. Wang, W. Zhong, Z. Wang, T. Yang, and X. Zhou, "G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation," https://omanscience.com/en/articles/g-2-ptq-improving-llm-post-training-quantization-with-generalized-gradient-compensation.