Abstract
Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibility between variable prefixes and fixed-shape graphs. Uniform truncation sacrifices acceptable tokens. We present DScale, preserving drafter architecture, weights, and full draft length. A separate 112K-parameter predictor requires neither confidence calibration nor hardware speed-curve preparation. Path-aware tiles reduce padding. Dynamic verify-length (DVL) allocation packs scored prefixes into half the native verification capacity. Fixed-address workspaces propagate changing boundaries through verification and acceptance while reusing captured graphs. On A100-40GB with tensor parallelism 1, Qwen3-8B and Qwen3-4B cover four datasets and concurrency 8-32, reusing each target's frozen predictor. Geometric-mean throughput gains across these configurations are respectively 43.9% and 48.8% over DFlash, 22.2% and 37.7% over DSpark, and 24.4% and 32.0% over Domino, with lower request latency. Cumulative ablations show that adding the three mechanisms successively increases geometric-mean throughput, while budget adjustment improves accepted-token retention. GPU profiling shows that complete decode-step time on GSM8K decreases by 30.8-52.5% relative to DFlash
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Chen, R., Xu, M., Fang, Z., Ye, K., & Xu, C. (2026). DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification. https://omanscience.com/en/articles/dscale-scaling-block-diffusion-speculative-decoding-with-adaptive-verification
MLA 9
Chen, Rongjian, et al. "DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification." https://omanscience.com/en/articles/dscale-scaling-block-diffusion-speculative-decoding-with-adaptive-verification.
Chicago (author–date)
Chen, Rongjian, Minxian Xu, Zhengxin Fang, Kejiang Ye, and Chengzhong Xu. 2026. "DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification." https://omanscience.com/en/articles/dscale-scaling-block-diffusion-speculative-decoding-with-adaptive-verification.
Harvard
Chen, R., Xu, M., Fang, Z., Ye, K. and Xu, C. (2026) 'DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification', Available at: https://omanscience.com/en/articles/dscale-scaling-block-diffusion-speculative-decoding-with-adaptive-verification.
Vancouver
Chen R, Xu M, Fang Z, Ye K, Xu C. DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification. https://omanscience.com/en/articles/dscale-scaling-block-diffusion-speculative-decoding-with-adaptive-verification
IEEE
R. Chen, M. Xu, Z. Fang, K. Ye, and C. Xu, "DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification," https://omanscience.com/en/articles/dscale-scaling-block-diffusion-speculative-decoding-with-adaptive-verification.