Preprint Open access
SlimKV: Joint Token-Feature KV Cache Compression with Reconstruction-Free Beacon Attention
Long-context LLM serving is increasingly bottlenecked by KV-cache memory, especially in resource-constrained scenarios. Among existing KV-cache compression strategies, token-wise methods reduce cached states but risk information loss through eviction or condensation, while feature-wise methods reduce per-token KV dimen …