Abstract

Block diffusion accelerates speculative decoding by drafting multiple tokens in one forward pass. However, each position predicts a marginal distribution without observing earlier proposed tokens, limiting draft quality and acceptance length. We identify a concrete failure, the \emph{repetition trap}, in which neighboring positions produce redundant copies of the same token. We explain this tendency theoretically and empirically examine its association with shorter accepted drafts. Recent methods refine marginal predictions with an additional causal head or a separately trained drafter, increasing parameter storage and introducing separate training objectives. We instead propose D-Loop, which introduces \emph{intra-block causal conditioning} within the original diffusion drafter without additional model components. Inspired by semi-autoregressive generation and parameter sharing, D-Loop reuses the same backbone across looped passes. The first pass proposes a block, and the second conditions on a selected prefix to regenerate the suffix in parallel. A complementary prefix--suffix objective trains the shared drafter for both anchor-only prefix prediction and prefix-conditioned suffix prediction. Across eight math, code, and chat benchmarks, D-Loop can beat DFlash and DSpark on Qwen3-4B and Qwen3-8B with obvious gains.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Chen, K., He, Y., Gong, C., Liu, H., Long, G., Li, J., Wu, S., Zhang, S., Li, H., Liu, Z., & Liu, R. (2026). D-Loop: Looped Diffusion Drafting for Speculative Decoding. https://omanscience.com/en/articles/d-loop-looped-diffusion-drafting-for-speculative-decoding

MLA 9

Chen, Kecheng, et al. "D-Loop: Looped Diffusion Drafting for Speculative Decoding." https://omanscience.com/en/articles/d-loop-looped-diffusion-drafting-for-speculative-decoding.

Chicago (author–date)

Chen, Kecheng, Yuyang He, Cheng Gong, Hui Liu, Guoping Long, Jiajun Li, Shi Wu, Suiyun Zhang, Haoliang Li, Ziru Liu, and Rui Liu. 2026. "D-Loop: Looped Diffusion Drafting for Speculative Decoding." https://omanscience.com/en/articles/d-loop-looped-diffusion-drafting-for-speculative-decoding.

Harvard

Chen, K., He, Y., Gong, C., Liu, H., Long, G., Li, J., Wu, S., Zhang, S., Li, H., Liu, Z. and Liu, R. (2026) 'D-Loop: Looped Diffusion Drafting for Speculative Decoding', Available at: https://omanscience.com/en/articles/d-loop-looped-diffusion-drafting-for-speculative-decoding.

Vancouver

Chen K, He Y, Gong C, Liu H, Long G, Li J, et al. D-Loop: Looped Diffusion Drafting for Speculative Decoding. https://omanscience.com/en/articles/d-loop-looped-diffusion-drafting-for-speculative-decoding

IEEE

K. Chen, Y. He, C. Gong, H. Liu, G. Long, J. Li, S. Wu, S. Zhang, H. Li, Z. Liu, and R. Liu, "D-Loop: Looped Diffusion Drafting for Speculative Decoding," https://omanscience.com/en/articles/d-loop-looped-diffusion-drafting-for-speculative-decoding.