الباحثون

Yuanhao Ban

المنشورات 3

نسخة أولية وصول مفتوح

Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards

Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference. In this work, we develop a simple and effective post-training recipe for open-domain text-to-image gen …

نسخة أولية وصول مفتوح

ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning

Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppresses under-generated …

المؤلفون المشاركون