Abstract
Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache. Low-bit quantization reduces storage and memory traffic, while query-channel pruning can further reduce key-cache reads. Rotation-based quantization redistributes the energy of key outliers across channels. To maintain computational invariance, the same orthogonal transform must be applied to queries, preserving query-key dot products. However, this rotation can disperse query energy, weakening the separation between a few large components to retain and many small ones to prune. We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict. Using calibrated query and key statistics, Dual-QK combines partial key whitening with a query-aligned basis to balance key scales for INT2 quantization and concentrate query energy for dynamic channel pruning. Channel-0 protection and bucket-relative RoPE support low-bit accuracy over long contexts. Experiments on four models across five generative benchmarks and long-context retrieval tasks show improved accuracy over OSCAR on most tasks at 40% query-channel sparsity. At a 128K context, Dual-QK provides $6.8\times$ KV-cache compression and an estimated $8.3\times$ reduction in KV read volume relative to unpruned BF16. Under the evaluated configurations, our SGLang implementation achieves up to $3.75\times$ the decoding throughput of unpruned BF16.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Whang, S., Oh, J., Kim, M., Seo, D., Shin, J., Kielian, G., Yoo, H. J., & Kim, S. (2026). Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches. https://omanscience.com/en/articles/dual-qk-sharp-queries-and-flat-keys-for-prunable-2-bit-kv-caches
MLA 9
Whang, Sunjoo, et al. "Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches." https://omanscience.com/en/articles/dual-qk-sharp-queries-and-flat-keys-for-prunable-2-bit-kv-caches.
Chicago (author–date)
Whang, Sunjoo, Jungjun Oh, Minsung Kim, Dongho Seo, Jisu Shin, Gregory Kielian, Hoi-Jun Yoo, and Sangjin Kim. 2026. "Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches." https://omanscience.com/en/articles/dual-qk-sharp-queries-and-flat-keys-for-prunable-2-bit-kv-caches.
Harvard
Whang, S., Oh, J., Kim, M., Seo, D., Shin, J., Kielian, G., Yoo, H. J. and Kim, S. (2026) 'Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches', Available at: https://omanscience.com/en/articles/dual-qk-sharp-queries-and-flat-keys-for-prunable-2-bit-kv-caches.
Vancouver
Whang S, Oh J, Kim M, Seo D, Shin J, Kielian G, et al. Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches. https://omanscience.com/en/articles/dual-qk-sharp-queries-and-flat-keys-for-prunable-2-bit-kv-caches
IEEE
S. Whang, J. Oh, M. Kim, D. Seo, J. Shin, G. Kielian, H. J. Yoo, and S. Kim, "Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches," https://omanscience.com/en/articles/dual-qk-sharp-queries-and-flat-keys-for-prunable-2-bit-kv-caches.