الباحثون

Haiying He

المنشورات 1

نسخة أولية وصول مفتوح

Frame Differential On-Policy Self-Distillation for Video Reasoning

Haiying He, Xin Zheng, Shaoli Hu وآخرون · 2026

Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly fine-grainedvisual or temporal credit assignment. In video reasoning, however, current RL methods typically train with a fixed sparse frame budget: increasing the number …

المؤلفون المشاركون