الملخص

Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction. A scalar reward cannot localize errors, causing optimization to perturb satisfactory tokens while under-targeting the tokens that actually need to change. We introduce Token-Level Video Reinforcement Learning, TVRL, a framework that derives token-level credit from the reward being optimized. Our key insight is that the answer likelihood of a frozen vision-language model provides both signals: its outputs contribute to the video-level reward, while magnitudes of its video-input gradients reveal which generated video tokens most affect that score. We instantiate TVRL in Group Relative Policy Optimization by averaging prompt-derived question rewards into one group-relative advantage and using detached, question-conditioned token-credit maps to reweight dense denoising-transition log-probabilities inside the clipped policy ratio. On VBench-2.0, TVRL achieves an Overall score of 57.69, outperforming the base model by 3.60 points. TVRL also improves matched GRPO baselines across three SDE samplers (SAGE, Flow, and Dance) by 2.68--3.15 points and across four reward models (VideoAlign, VideoScore2, UnifiedReward2, and Qwen3.5-9B) by 1.33--3.15 points.

الكلمات المفتاحية

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Wang, Y., Qian, G. G., Li, Y., Kag, A., & Fu, Y. (2026). Token-Level Video Reinforcement Learning. https://omanscience.com/ar/articles/token-level-video-reinforcement-learning

MLA 9

Wang, Yifan, et al. "Token-Level Video Reinforcement Learning." https://omanscience.com/ar/articles/token-level-video-reinforcement-learning.

شيكاغو (المؤلف–التاريخ)

Wang, Yifan, Gordon Guocheng Qian, Yanyu Li, Anil Kag, and Yun Fu. 2026. "Token-Level Video Reinforcement Learning." https://omanscience.com/ar/articles/token-level-video-reinforcement-learning.

هارفارد

Wang, Y., Qian, G. G., Li, Y., Kag, A. and Fu, Y. (2026) 'Token-Level Video Reinforcement Learning', Available at: https://omanscience.com/ar/articles/token-level-video-reinforcement-learning.

فانكوفر

Wang Y, Qian GG, Li Y, Kag A, Fu Y. Token-Level Video Reinforcement Learning. https://omanscience.com/ar/articles/token-level-video-reinforcement-learning

IEEE

Y. Wang, G. G. Qian, Y. Li, A. Kag, and Y. Fu, "Token-Level Video Reinforcement Learning," https://omanscience.com/ar/articles/token-level-video-reinforcement-learning.