Abstract
Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at https://github.com/naver-ai/DLoop.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Gu, G., Heo, B., Jun, H., Kang, Y., Lee, S., Yun, S., & Han, D. (2026). DLoop: Looped Speculative Decoding. https://omanscience.com/en/articles/dloop-looped-speculative-decoding
MLA 9
Gu, Geonmo, et al. "DLoop: Looped Speculative Decoding." https://omanscience.com/en/articles/dloop-looped-speculative-decoding.
Chicago (author–date)
Gu, Geonmo, Byeongho Heo, HeeJae Jun, Yoohoon Kang, Sangmin Lee, Sangdoo Yun, and Dongyoon Han. 2026. "DLoop: Looped Speculative Decoding." https://omanscience.com/en/articles/dloop-looped-speculative-decoding.
Harvard
Gu, G., Heo, B., Jun, H., Kang, Y., Lee, S., Yun, S. and Han, D. (2026) 'DLoop: Looped Speculative Decoding', Available at: https://omanscience.com/en/articles/dloop-looped-speculative-decoding.
Vancouver
Gu G, Heo B, Jun H, Kang Y, Lee S, Yun S, et al. DLoop: Looped Speculative Decoding. https://omanscience.com/en/articles/dloop-looped-speculative-decoding
IEEE
G. Gu, B. Heo, H. Jun, Y. Kang, S. Lee, S. Yun, and D. Han, "DLoop: Looped Speculative Decoding," https://omanscience.com/en/articles/dloop-looped-speculative-decoding.