Abstract
Masked diffusion language models (dLLMs) have shown strong potential for faster inference through parallel token generation when combined with confidence-based samplers. However, recent work has shown that such methods can defer unmasking high-entropy fork positions at which multiple plausible continuations exist. This results in reduced generation diversity, as shown by worse pass@k scaling, and limits gains obtainable from RL post-training. To avoid this flexibility trap, prior work advocated for autoregressive (AR) sampling. Here, we show that discarding confidence-based sampling is unnecessary and, once inference cost is taken into account, wasteful. We first propose Fork-dLLM, a simple hybrid sampler that uses AR-style ordering only at uncertain fallback steps while retaining parallel generation otherwise. We then extend the same principle to post-training with ForkGRPO, which uses Fork-dLLM rollouts and applies the GRPO objective only at fallback steps, preserving exact policy-likelihood ratios while substantially reducing rollout and optimization cost. In our experiments, Fork-dLLM matches the strong pass@k scaling of AR sampling while being 2-3x more efficient, and ForkGRPO achieves downstream performance comparable to or better than AR-based GRPO baselines at a substantially lower training cost.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Frković, S., Jazbec, M., & Naesseth, C. A. (2026). Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models. https://omanscience.com/en/articles/fork-dllm-avoiding-the-flexibility-trap-in-diffusion-language-models
MLA 9
Frković, Stipe, et al. "Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models." https://omanscience.com/en/articles/fork-dllm-avoiding-the-flexibility-trap-in-diffusion-language-models.
Chicago (author–date)
Frković, Stipe, Metod Jazbec, and Christian A. Naesseth. 2026. "Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models." https://omanscience.com/en/articles/fork-dllm-avoiding-the-flexibility-trap-in-diffusion-language-models.
Harvard
Frković, S., Jazbec, M. and Naesseth, C. A. (2026) 'Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models', Available at: https://omanscience.com/en/articles/fork-dllm-avoiding-the-flexibility-trap-in-diffusion-language-models.
Vancouver
Frković S, Jazbec M, Naesseth CA. Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models. https://omanscience.com/en/articles/fork-dllm-avoiding-the-flexibility-trap-in-diffusion-language-models
IEEE
S. Frković, M. Jazbec, and C. A. Naesseth, "Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models," https://omanscience.com/en/articles/fork-dllm-avoiding-the-flexibility-trap-in-diffusion-language-models.