Authors

Xiao-Wen Chang

Publications 3

Preprint Open access

JARQ: Joint Alternating Refinement for Quantization

Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input correlations couple the errors of different groups, and useful code changes often involve ma …

Co-authors