Abstract

As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textit{heterogeneous} per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \textit{real-world} serving experiments, DWS reduces cost-prediction error by up to $2.50\times$ over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to $1.92\times$ for online chatbots.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zheng, H., Fu, F., Yuan, B., Zhang, Y., Deng, L., Wang, H., Zhu, Y., Yan, X., & Jiang, J. (2026). Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving. https://omanscience.com/en/articles/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm-serving

MLA 9

Zheng, Haoyu, et al. "Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving." https://omanscience.com/en/articles/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm-serving.

Chicago (author–date)

Zheng, Haoyu, Fangcheng Fu, Binhang Yuan, Yongqiang Zhang, Liang Deng, Hao Wang, Yuanyuan Zhu, Xiao Yan, and Jiawei Jiang. 2026. "Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving." https://omanscience.com/en/articles/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm-serving.

Harvard

Zheng, H., Fu, F., Yuan, B., Zhang, Y., Deng, L., Wang, H., Zhu, Y., Yan, X. and Jiang, J. (2026) 'Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving', Available at: https://omanscience.com/en/articles/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm-serving.

Vancouver

Zheng H, Fu F, Yuan B, Zhang Y, Deng L, Wang H, et al. Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving. https://omanscience.com/en/articles/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm-serving

IEEE

H. Zheng, F. Fu, B. Yuan, Y. Zhang, L. Deng, H. Wang, Y. Zhu, X. Yan, and J. Jiang, "Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving," https://omanscience.com/en/articles/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm-serving.