الملخص

Diffusion large language models (dLLMs) offer a promising parallel decoding paradigm as an alternative to autoregressive generation through iterative unmasking. However, dLLMs typically require many steps before token confidence reaches the decoding threshold, resulting in inefficient inference even with block-wise KV caching. To accelerate dLLM inference, we for the first time propose an "early-bird (EB)" decoding framework, motivated by the observation that tokens with similarly low entropy tend to cluster and can be jointly decoded earlier, before reaching the confidence threshold. In particular, our EB-Decode framework integrates two key enablers: (1) a learnable network that adaptively groups tokens with similar uncertainty into variable-length blocks, rather than relying on fixed block sizes; (2) a position-aware sampler that learns to unmask tokens in parallel using fewer decoding steps within predicted variable-length blocks. Both components are developed without modifying pretrained dLLM weights and can therefore be directly deployed as plug-ins during serving, with negligible training and inference overhead. Extensive experiments across three models and four benchmarks consistently validate our observation and the effectiveness of EB-Decode, achieving 3.53-18.76$\times$ higher throughput than the vanilla decoding method and up to 1.58$\times$ higher throughput over the strongest baseline, Fast-dLLM, with comparable accuracy.

الكلمات المفتاحية

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Wei, L., Zhou, W., Wu, J., Shen, Y., Wang, M., & You, H. (2026). Early-Bird Decoding: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling. https://omanscience.com/ar/articles/early-bird-decoding-accelerating-diffusion-llms-with-learnable-block-sizes-and-parallel-sampling

MLA 9

Wei, Lixuan, et al. "Early-Bird Decoding: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling." https://omanscience.com/ar/articles/early-bird-decoding-accelerating-diffusion-llms-with-learnable-block-sizes-and-parallel-sampling.

شيكاغو (المؤلف–التاريخ)

Wei, Lixuan, Wei Zhou, Jianwen Wu, Yipeng Shen, Meiling Wang, and Haoran You. 2026. "Early-Bird Decoding: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling." https://omanscience.com/ar/articles/early-bird-decoding-accelerating-diffusion-llms-with-learnable-block-sizes-and-parallel-sampling.

هارفارد

Wei, L., Zhou, W., Wu, J., Shen, Y., Wang, M. and You, H. (2026) 'Early-Bird Decoding: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling', Available at: https://omanscience.com/ar/articles/early-bird-decoding-accelerating-diffusion-llms-with-learnable-block-sizes-and-parallel-sampling.

فانكوفر

Wei L, Zhou W, Wu J, Shen Y, Wang M, You H. Early-Bird Decoding: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling. https://omanscience.com/ar/articles/early-bird-decoding-accelerating-diffusion-llms-with-learnable-block-sizes-and-parallel-sampling

IEEE

L. Wei, W. Zhou, J. Wu, Y. Shen, M. Wang, and H. You, "Early-Bird Decoding: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling," https://omanscience.com/ar/articles/early-bird-decoding-accelerating-diffusion-llms-with-learnable-block-sizes-and-parallel-sampling.