نسخة أولية وصول مفتوح
Activation Denoising: A Robustness View on Parallel vs Sequential LLM Quantization
Post-training quantization is a powerful tool for compressing large language models. The most scalable methods quantize every layer in parallel, but quantization errors then compound through the residual stream, as no layer corrects for the errors of the layers before it. Sequential quantization accounts for this error …