Authors

Gordon Guocheng Qian

Publications 3

Preprint Open access

Token-Level Video Reinforcement Learning

Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction. A scalar reward cannot localize errors, causing optimization to perturb satisfactory t …

Co-authors