Preprint Open access
DRelay: Global Draft Context for Prefix-Aware Parallel Speculative Decoding Repair
Parallel drafting reduces the drafting overhead of speculative decoding for large language models (LLMs), but its gains remain limited by the accepted prefix length. Even when the correct token is present in the candidate pool, a single early selection error prevents subsequent predictions from being used. We propose D …