الملخص

Diffusion-based LLMs (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs by enabling bidirectional parallel refinement, alleviating the sequential decoding bottleneck of AR generation. However, their parallel iterative refinement mismatches AR accelerators optimized for sequential decoding and their discrete token generation differs from DiT accelerators designed for continuous denoising. Recent dLLM accelerators have explored workload-specific optimizations to reduce vocabulary processing overhead and redundant computation across denoising iterations. However, these approaches retain all tokens in parallel execution, despite varying token refinement utility and execution requirements. This paper presents DynaTE, a hardware--software co-design architecture that dynamically adapts accelerator execution to evolving token states during dLLM decoding. DynaTE first enables adaptive token execution by skipping low-utility token computation, while a dimension-reconfigurable PE array maintains high utilization under varying active-token patterns. Second, DynaTE exploits dynamic token dependencies through FLDD to refine a small number of locally dependent tokens within the current iteration, reducing the overall number of denoising iterations, while a Merge--Split--Merge dataflow hides the resulting serial overhead. Third, a streaming vocabulary engine interleaves multiple token streams from the LM head to accommodate irregular output variations caused by selective token computation and uneven vocabulary-selection demands. Evaluated on two representative dLLMs, DynaTE achieves 2.05--2.78$\times$ speedup and 2.99--3.93$\times$ higher energy efficiency over state-of-the-art dLLM accelerators, while delivering 2.55$\times$ speedup and 6.07$\times$ higher energy efficiency over Jetson AGX Orin.

الكلمات المفتاحية

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Jiang, M., Wang, J., Li, S., Shen, H., & Huang, K. (2026). DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution. https://omanscience.com/ar/articles/dynate-accelerating-diffusion-llms-via-dynamic-token-execution

MLA 9

Jiang, Minghan, et al. "DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution." https://omanscience.com/ar/articles/dynate-accelerating-diffusion-llms-via-dynamic-token-execution.

شيكاغو (المؤلف–التاريخ)

Jiang, Minghan, Jiayi Wang, Shuaiting Li, Haibin Shen, and Kejie Huang. 2026. "DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution." https://omanscience.com/ar/articles/dynate-accelerating-diffusion-llms-via-dynamic-token-execution.

هارفارد

Jiang, M., Wang, J., Li, S., Shen, H. and Huang, K. (2026) 'DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution', Available at: https://omanscience.com/ar/articles/dynate-accelerating-diffusion-llms-via-dynamic-token-execution.

فانكوفر

Jiang M, Wang J, Li S, Shen H, Huang K. DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution. https://omanscience.com/ar/articles/dynate-accelerating-diffusion-llms-via-dynamic-token-execution

IEEE

M. Jiang, J. Wang, S. Li, H. Shen, and K. Huang, "DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution," https://omanscience.com/ar/articles/dynate-accelerating-diffusion-llms-via-dynamic-token-execution.