الملخص

The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade substantially in the 1-bit regime. At such extreme compression, each codebook must represent a larger group of channels with a limited set of centroids, making effective use of its capacity increasingly challenging. To address this, we introduce $\textbf{TaSQ}$, which tailors the VQ target space by combining query-guided channel weighting, cross-head normalization, and covariance-aware channel grouping to better reflect the error sensitivity and statistical structure of cached activations. Since these transforms are RoPE-compatible and can be easily merged into projection weights and codebooks, TaSQ preserves the conventional VQ lookup structure and adds negligible serving overhead. Across general, long-chain-of-thought reasoning, and long-context retrieval benchmarks, TaSQ consistently outperforms existing low-bit KV cache VQ baselines while preserving reasoning stability. On a single RTX 6000 Ada GPU, its SGLang implementation supports up to $14\times$ larger batch sizes and achieves $1.87\times$ higher peak throughput compared to the BF16 baseline.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Cheong, M., Son, D., & Yoo, S. (2026). Tailoring the Quantization Space for 1-Bit KV Cache Compression. https://omanscience.com/ar/articles/tailoring-the-quantization-space-for-1-bit-kv-cache-compression

MLA 9

Cheong, Minsoo, et al. "Tailoring the Quantization Space for 1-Bit KV Cache Compression." https://omanscience.com/ar/articles/tailoring-the-quantization-space-for-1-bit-kv-cache-compression.

شيكاغو (المؤلف–التاريخ)

Cheong, Minsoo, Donghyun Son, and Sungjoo Yoo. 2026. "Tailoring the Quantization Space for 1-Bit KV Cache Compression." https://omanscience.com/ar/articles/tailoring-the-quantization-space-for-1-bit-kv-cache-compression.

هارفارد

Cheong, M., Son, D. and Yoo, S. (2026) 'Tailoring the Quantization Space for 1-Bit KV Cache Compression', Available at: https://omanscience.com/ar/articles/tailoring-the-quantization-space-for-1-bit-kv-cache-compression.

فانكوفر

Cheong M, Son D, Yoo S. Tailoring the Quantization Space for 1-Bit KV Cache Compression. https://omanscience.com/ar/articles/tailoring-the-quantization-space-for-1-bit-kv-cache-compression

IEEE

M. Cheong, D. Son, and S. Yoo, "Tailoring the Quantization Space for 1-Bit KV Cache Compression," https://omanscience.com/ar/articles/tailoring-the-quantization-space-for-1-bit-kv-cache-compression.