الملخص

Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. We propose stepwise risk-sensitive GRPO (StepRS-GRPO), which varies the risk coefficient of the group-advantage transformation across denoising states while retaining the underlying trainer. For binary rewards, we show that this transformation is exactly a prompt- and state-dependent rescaling of centered outcome advantages. A capability-based calibration suggests a coefficient scale, while endpoint and interpolation ablations guide schedule selection. Across multiple dLLM backbones and mathematical reasoning benchmarks, StepRS-GRPO improves both pass@1 accuracy and pass@k coverage over centered GRPO, while increasing answer diversity. In our ablation studies, mass-matched controls support the contributions of state allocation and schedule direction, and the gains persist after matching the root mean square (RMS) of the advantages to that of centered GRPO. Reasoning-trace diagnostics further show that the diversity gains from StepRS-GRPO extend beyond final-answer strings.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Yu, Y., Zuo, B., Crandall, D., Zhu, Y., & Zhou, D. (2026). Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models. https://omanscience.com/ar/articles/exploring-more-reasoning-better-stepwise-risk-sensitive-grpo-for-diffusion-language-models

MLA 9

Yu, Yue, et al. "Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models." https://omanscience.com/ar/articles/exploring-more-reasoning-better-stepwise-risk-sensitive-grpo-for-diffusion-language-models.

شيكاغو (المؤلف–التاريخ)

Yu, Yue, Bowen Zuo, David Crandall, Yinglun Zhu, and Dongruo Zhou. 2026. "Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models." https://omanscience.com/ar/articles/exploring-more-reasoning-better-stepwise-risk-sensitive-grpo-for-diffusion-language-models.

هارفارد

Yu, Y., Zuo, B., Crandall, D., Zhu, Y. and Zhou, D. (2026) 'Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models', Available at: https://omanscience.com/ar/articles/exploring-more-reasoning-better-stepwise-risk-sensitive-grpo-for-diffusion-language-models.

فانكوفر

Yu Y, Zuo B, Crandall D, Zhu Y, Zhou D. Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models. https://omanscience.com/ar/articles/exploring-more-reasoning-better-stepwise-risk-sensitive-grpo-for-diffusion-language-models

IEEE

Y. Yu, B. Zuo, D. Crandall, Y. Zhu, and D. Zhou, "Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models," https://omanscience.com/ar/articles/exploring-more-reasoning-better-stepwise-risk-sensitive-grpo-for-diffusion-language-models.