الباحثون

Chak Tou Leong

المنشورات 1

نسخة أولية وصول مفتوح

ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation

Xinghao Chen, Junnan Dong, Cai Ke وآخرون · 2026

Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and reusing its key-value …

المؤلفون المشاركون