Abstract
Long-running agents repeatedly call an LLM while retaining most of their document window, evicting old documents, and appending new ones. These rolling updates break exact prefix caching and motivate non-prefix KV-cache reuse with selective recomputation. We show that persistent KV-cache reuse with selective recomputation can be history-dependent: in our rolling-agent workload, an unchanged prompt can produce different answers depending on the requests processed before it. At a matched 5% recomputation budget, document-aligned recomputation reduces answer variation across request orders from 69.0% with CacheBlend's token top-$k$ policy to 26.1%. When each prompt is evaluated after a different sequence of preceding requests, document-aligned recomputation improves fidelity to full prefill by 34.5-52.5 percentage points over token top-$k$, while both policies achieve approximately 5.7$\times$ median TTFT speedup. Our ablation study shows that, in our rolling-agent workload, contiguity is the main factor associated with robust selective recomputation.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Gu, T., Guan, A., Huang, M., Rao, N., Shah, S., & Capetz, M. (2026). Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents. https://omanscience.com/en/articles/request-order-matters-cache-history-sensitivity-in-selective-kv-cache-reuse-for-rolling-agents
MLA 9
Gu, Tiffany, et al. "Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents." https://omanscience.com/en/articles/request-order-matters-cache-history-sensitivity-in-selective-kv-cache-reuse-for-rolling-agents.
Chicago (author–date)
Gu, Tiffany, Annie Guan, Manshu Huang, Nitin Rao, Siddhant Shah, and Margaret Capetz. 2026. "Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents." https://omanscience.com/en/articles/request-order-matters-cache-history-sensitivity-in-selective-kv-cache-reuse-for-rolling-agents.
Harvard
Gu, T., Guan, A., Huang, M., Rao, N., Shah, S. and Capetz, M. (2026) 'Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents', Available at: https://omanscience.com/en/articles/request-order-matters-cache-history-sensitivity-in-selective-kv-cache-reuse-for-rolling-agents.
Vancouver
Gu T, Guan A, Huang M, Rao N, Shah S, Capetz M. Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents. https://omanscience.com/en/articles/request-order-matters-cache-history-sensitivity-in-selective-kv-cache-reuse-for-rolling-agents
IEEE
T. Gu, A. Guan, M. Huang, N. Rao, S. Shah, and M. Capetz, "Request Order Matters: Cache-History Sensitivity in Selective KV-Cache Reuse for Rolling Agents," https://omanscience.com/en/articles/request-order-matters-cache-history-sensitivity-in-selective-kv-cache-reuse-for-rolling-agents.