نسخة أولية وصول مفتوح
EchoPress: Query-Agnostic KV Cache Pruning via Virtual Context Reconstruction
KV cache pruning reduces long-context inference memory usage by evicting less important key-value pairs. KVzip estimates importance through context reconstruction: prompting a model to repeat the context chunk by chunk. This achieves strong compression quality at the cost of additional forward passes. Learned approxima …