نسخة أولية وصول مفتوح
Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents
Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones. These rolling updates break exact prefix caching and motivate non-prefix KV-cache reuse with selective recomputation. We show that persistent KV-cache reuse with selective recomputat …