Preprint Open access
Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration
KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contrib …