نسخة أولية وصول مفتوح
BARQ: Balanced Codebook Refinement for Low-Bit LLM Quantization
As large language models (LLMs) grow in parameter count, model storage and parameter memory traffic have become major bottlenecks to efficient deployment. Codebook-based weight quantization reduces these costs, but imbalanced nearest-codeword assignments during fitting can leave some codewords insufficiently updated, l …