نسخة أولية وصول مفتوح
SGA-Flow-GRPO: Spatial Gradient-Guided Credit Assignment for Flow-GRPO
Reinforcement Learning (RL) has proven effective in aligning flow-based generative models with human preferences. Recently, Flow-GRPO has emerged as an efficient critic-free paradigm by calculating advantages over sampled candidate trajectories. However, standard Flow-GRPO applies a uniform scalar advantage across both …