[
    {
        "id": "osp-21388",
        "type": "article-journal",
        "title": "Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving",
        "author": [
            {
                "family": "Zheng",
                "given": "Haoyu"
            },
            {
                "family": "Fu",
                "given": "Fangcheng"
            },
            {
                "family": "Yuan",
                "given": "Binhang"
            },
            {
                "family": "Zhang",
                "given": "Yongqiang"
            },
            {
                "family": "Deng",
                "given": "Liang"
            },
            {
                "family": "Wang",
                "given": "Hao"
            },
            {
                "family": "Zhu",
                "given": "Yuanyuan"
            },
            {
                "family": "Yan",
                "given": "Xiao"
            },
            {
                "family": "Jiang",
                "given": "Jiawei"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/denoising-surface-modeling-and-predicting-inference-cost-for-diffusion-llm-serving",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \\textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \\textit{heterogeneous} per-step costs. We observe that the block-autoregressive generation mechanism induces a two-dimensional execution structure over output blocks and within-block denoising steps, whereas these proxies collapse it into a scalar, discarding information essential for characterizing the cost. Motivated by this insight, we propose the Denoising Workload Surface (DWS), which preserves this two-dimensional block-step structure as a probability surface to weight the heterogeneous per-step costs. We then design a coarse-to-fine training scheme that enables a lightweight prompt-only predictor to accurately predict the complex DWS. This predictor runs efficiently even on a single CPU core, avoiding GPU contention with the serving model. Since DWS decouples request-dependent execution behavior from deployment-specific cost factors, the predictor transfers across hardware configurations without retraining. In \\textit{real-world} serving experiments, DWS reduces cost-prediction error by up to $2.50\\times$ over scalar-based predictors, while the DWS-guided shortest-job-first scheduler reduces end-to-end latency by up to $1.92\\times$ for online chatbots."
    }
]