Preprint Open access
SharpDraft: Accelerating Long-Context Speculative Decoding with Cardinality-Aware Query Scaling
Long-form reasoning makes inference expensive, and speculative decoding mitigates this cost by verifying multiple draft tokens in parallel. Its speedup, however, can fade as context grows and draft acceptance declines. We focus on attention-mass dilution: as softmax normalizes over more visible Keys, the mass concentra …