Preprint Open access
Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference
The memory usage and decoding latency of LLM inference grow rapidly with context length. To reduce these costs, key-value (KV) cache compression methods selectively retain cached states based on token importance or differences in attention patterns across heads. However, we discover that retrieval capability varies sub …