نسخة أولية وصول مفتوح
WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds
Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. Across three models (1.5B-8B) and three regimes (needle retrieval, long chain-of-thought, and multi-turn recall), we measure head behavior on four model-regime com …