Preprint Open access
Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches
Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache. Low-bit quantization reduces storage and memory traffic, while query-channel pruning can further reduce key-cache reads. Rotation-based quantization redistributes the energy of key outliers across channels. To maintai …