الباحثون

Guoping Long

المنشورات 3

نسخة أولية وصول مفتوح

QUILT: Rethinking Sparse-Attention Prefill through Shared Query Execution

Zhenduo Zhao, Qihui Zhou, Mingcong Song وآخرون · 2026

Sparse attention reduces the cost of long-context attention, but existing kernels typically process queries independently, repeatedly loading and dequantizing KV entries shared across queries. We observe substantial overlap in the KV entries selected by neighboring queries, creating opportunities for cross-query reuse. …

نسخة أولية وصول مفتوح

D-Loop: Looped Diffusion Drafting for Speculative Decoding

Kecheng Chen, Yuyang He, Cheng Gong وآخرون · 2026

Block diffusion accelerates speculative decoding by drafting multiple tokens in one forward pass. However, each position predicts a marginal distribution without observing earlier proposed tokens, limiting draft quality and acceptance length. We identify a concrete failure, the \emph{repetition trap}, in which neighbor …

نسخة أولية وصول مفتوح

From Experience to Expertise: Adoption-Aware Memory Learning for Data-Scarce NPU Kernel Synthesis

Longxiao Fan, Tao Zhang, Han Yan وآخرون · 2026

High-performance kernels underpin efficient accelerator execution but require expert tuning and lengthy manual optimization cycles. LLM coding agents promise automation, yet their CUDA knowledge transfers poorly to data-scarce domain-specific architectures (DSAs) such as NPUs, whose execution models and memory hierarch …

المؤلفون المشاركون