الباحثون

Haocheng Sun

المنشورات 2

نسخة أولية وصول مفتوح

ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation

Xinghao Chen, Junnan Dong, Cai Ke وآخرون · 2026

Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and reusing its key-value …

نسخة أولية وصول مفتوح

Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding

Haohui Zhang, Keyu Chen, Haocheng Sun وآخرون · 2026

Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor i …

المؤلفون المشاركون