الباحثون

Seongwoon Jo

المنشورات 1

نسخة أولية وصول مفتوح

SharpDraft: Accelerating Long-Context Speculative Decoding with Cardinality-Aware Query Scaling

Jeonghoon Park, Seongwoon Jo, Jongwon Lee وآخرون · 2026

Long-form reasoning makes inference expensive, and speculative decoding mitigates this cost by verifying multiple draft tokens in parallel. Its speedup, however, can fade as context grows and draft acceptance declines. We focus on attention-mass dilution: as softmax normalizes over more visible Keys, the mass concentra …

المؤلفون المشاركون