الملخص
Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This noise decides which rollouts the model learns from, yet it perturbs every channel and spatial position of the latent equally. In this paper, we instead show that latent elements differ in how much they change the generated image, so exploration should adapt to these differences. We introduce EXPLORENET to learn an adaptive exploration distribution. EXPLORENET is a policy that predicts a noise scale for every latent element from the current latent, the denoising step, and the prompt, before any reward is observed; it is trained on the reward spread of each rollout group and discarded after training, leaving inference unchanged. On Stable Diffusion 3.5 Medium, EXPLORENET improves held-out GenEval2 by 14% over Flow-GRPO, transfers to two independent compositional benchmarks and five preference and image-quality models, and reaches a 67.2% human preference win-rate. Overall, across our group-relative diffusion RL experiments, we find that exploration is learnable, the shape of the exploration distribution outweighs its magnitude, and rollout quality is more effective than rollout quantity.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Li, S. S., Han, X., Tsvetkov, Y., & Zettlemoyer, L. (2026). ExploreNet: Learning Where to Explore in Diffusion GRPO. https://omanscience.com/ar/articles/explorenet-learning-where-to-explore-in-diffusion-grpo
MLA 9
Li, Shuyue Stella, et al. "ExploreNet: Learning Where to Explore in Diffusion GRPO." https://omanscience.com/ar/articles/explorenet-learning-where-to-explore-in-diffusion-grpo.
شيكاغو (المؤلف–التاريخ)
Li, Shuyue Stella, Xiaochuang Han, Yulia Tsvetkov, and Luke Zettlemoyer. 2026. "ExploreNet: Learning Where to Explore in Diffusion GRPO." https://omanscience.com/ar/articles/explorenet-learning-where-to-explore-in-diffusion-grpo.
هارفارد
Li, S. S., Han, X., Tsvetkov, Y. and Zettlemoyer, L. (2026) 'ExploreNet: Learning Where to Explore in Diffusion GRPO', Available at: https://omanscience.com/ar/articles/explorenet-learning-where-to-explore-in-diffusion-grpo.
فانكوفر
Li SS, Han X, Tsvetkov Y, Zettlemoyer L. ExploreNet: Learning Where to Explore in Diffusion GRPO. https://omanscience.com/ar/articles/explorenet-learning-where-to-explore-in-diffusion-grpo
IEEE
S. S. Li, X. Han, Y. Tsvetkov, and L. Zettlemoyer, "ExploreNet: Learning Where to Explore in Diffusion GRPO," https://omanscience.com/ar/articles/explorenet-learning-where-to-explore-in-diffusion-grpo.