Abstract

Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy across denoising steps, we identify a complementary redundancy axis within each speculative verification step: multi-branch computational redundancy. During speculative verification, draft branches inherit most tokens from their parents while unmasking a small set of additional positions, causing large portions of hidden states to remain highly similar across branches. We propose SpecFold, an algorithm-system co-design that exploits this multi-branch redundancy to reduce the cost of multi-branch speculative verification. Algorithmically, SpecFold performs token-level residual gating and selectively reuses parent computation through folded attention and FFN while preserving residual hidden states. Systemically, a Triton kernel implementation translates this fine-grained reuse into end-to-end throughput gains through efficient sparse multi-branch execution. SpecFold is orthogonal to temporal caching and compatible with existing DLLM speculation strategies. Across two DLLM families, five models, and five standard benchmarks, SpecFold achieves up to 1.64x throughput over Spiffy and up to 1.99x over vanilla decoding, while maintaining comparable task performance.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Ho, C. E., Sun, W., Shih, C. J., Li, H., Liu, Y., & Lin, Y. C. (2026). SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models. https://omanscience.com/en/articles/specfold-folding-multi-branch-redundancy-for-faster-speculative-decoding-in-diffusion-language-models

MLA 9

Ho, Chung-En, et al. "SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models." https://omanscience.com/en/articles/specfold-folding-multi-branch-redundancy-for-faster-speculative-decoding-in-diffusion-language-models.

Chicago (author–date)

Ho, Chung-En, Weiyu Sun, Cheng-Jhih Shih, He Li, Yong Liu, and Yingyan Celine Lin. 2026. "SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models." https://omanscience.com/en/articles/specfold-folding-multi-branch-redundancy-for-faster-speculative-decoding-in-diffusion-language-models.

Harvard

Ho, C. E., Sun, W., Shih, C. J., Li, H., Liu, Y. and Lin, Y. C. (2026) 'SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models', Available at: https://omanscience.com/en/articles/specfold-folding-multi-branch-redundancy-for-faster-speculative-decoding-in-diffusion-language-models.

Vancouver

Ho CE, Sun W, Shih CJ, Li H, Liu Y, Lin YC. SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models. https://omanscience.com/en/articles/specfold-folding-multi-branch-redundancy-for-faster-speculative-decoding-in-diffusion-language-models

IEEE

C. E. Ho, W. Sun, C. J. Shih, H. Li, Y. Liu, and Y. C. Lin, "SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models," https://omanscience.com/en/articles/specfold-folding-multi-branch-redundancy-for-faster-speculative-decoding-in-diffusion-language-models.