الملخص

Long-form reasoning makes inference expensive, and speculative decoding mitigates this cost by verifying multiple draft tokens in parallel. Its speedup, however, can fade as context grows and draft acceptance declines. We focus on attention-mass dilution: as softmax normalizes over more visible Keys, the mass concentrated on the highest-scoring Keys can decrease. We introduce SharpDraft, a training-free method that counteracts this effect through cardinality-aware Query scaling, without the computational overhead of online adaptation. Under explicit assumptions, we derive an exact top-$k$ mass correction and deploy a closed-form fixed-slope approximation. Across AIME-26, GPQA-Diamond, and LongGenBench Diary, SharpDraft achieves $2.59$-$3.19\times$ geometric-mean end-to-end speedups over target-only autoregressive decoding when applied to DFlash, PARD, and EAGLE 3.1. With DFlash, it improves decoding speed and outperforms full-parameter and LoRA-based online adaptation in end-to-end speedup, while matching the unmodified drafter's reported peak allocated GPU memory.

الكلمات المفتاحية

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Park, J., Jo, S., Lee, J., & Gong, T. (2026). SharpDraft: Accelerating Long-Context Speculative Decoding with Cardinality-Aware Query Scaling. https://omanscience.com/ar/articles/sharpdraft-accelerating-long-context-speculative-decoding-with-cardinality-aware-query-scaling

MLA 9

Park, Jeonghoon, et al. "SharpDraft: Accelerating Long-Context Speculative Decoding with Cardinality-Aware Query Scaling." https://omanscience.com/ar/articles/sharpdraft-accelerating-long-context-speculative-decoding-with-cardinality-aware-query-scaling.

شيكاغو (المؤلف–التاريخ)

Park, Jeonghoon, Seongwoon Jo, Jongwon Lee, and Taesik Gong. 2026. "SharpDraft: Accelerating Long-Context Speculative Decoding with Cardinality-Aware Query Scaling." https://omanscience.com/ar/articles/sharpdraft-accelerating-long-context-speculative-decoding-with-cardinality-aware-query-scaling.

هارفارد

Park, J., Jo, S., Lee, J. and Gong, T. (2026) 'SharpDraft: Accelerating Long-Context Speculative Decoding with Cardinality-Aware Query Scaling', Available at: https://omanscience.com/ar/articles/sharpdraft-accelerating-long-context-speculative-decoding-with-cardinality-aware-query-scaling.

فانكوفر

Park J, Jo S, Lee J, Gong T. SharpDraft: Accelerating Long-Context Speculative Decoding with Cardinality-Aware Query Scaling. https://omanscience.com/ar/articles/sharpdraft-accelerating-long-context-speculative-decoding-with-cardinality-aware-query-scaling

IEEE

J. Park, S. Jo, J. Lee, and T. Gong, "SharpDraft: Accelerating Long-Context Speculative Decoding with Cardinality-Aware Query Scaling," https://omanscience.com/ar/articles/sharpdraft-accelerating-long-context-speculative-decoding-with-cardinality-aware-query-scaling.