الملخص
Video-language models commonly assume that more temporal observations lead to more reliable reasoning. We question this assumption and argue that the key challenge is not merely processing more video frames efficiently, but learning to reason reliably under limited temporal evidence. We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference. During reinforcement post-training, paired dense and sparse views are optimized with grounding rewards and a reliability-gated reference reward, encouraging sparse view predictions to preserve task-relevant temporal evidence. Notably, SAVER is trained only on 1,250 randomly sampled temporal grounding examples, without using any video question answering annotations. Across three temporal grounding benchmarks and six video question-answering benchmarks, SAVER consistently improves performance across frame budgets. In particular, SAVER can match or surpass dense-frame Qwen3.5 baselines while using substantially fewer frames. These results show that temporal grounding can serve as an effective evidence-localization proxy for learning sparse video reasoning that transfers to broader video understanding tasks.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Sun, W., Du, Y., & Snoek, C. G. M. (2026). Less from More: Reinforcing Sparse Video Reasoning from Dense References. https://omanscience.com/ar/articles/less-from-more-reinforcing-sparse-video-reasoning-from-dense-references
MLA 9
Sun, Wenfang, et al. "Less from More: Reinforcing Sparse Video Reasoning from Dense References." https://omanscience.com/ar/articles/less-from-more-reinforcing-sparse-video-reasoning-from-dense-references.
شيكاغو (المؤلف–التاريخ)
Sun, Wenfang, Yingjun Du, and Cees G. M. Snoek. 2026. "Less from More: Reinforcing Sparse Video Reasoning from Dense References." https://omanscience.com/ar/articles/less-from-more-reinforcing-sparse-video-reasoning-from-dense-references.
هارفارد
Sun, W., Du, Y. and Snoek, C. G. M. (2026) 'Less from More: Reinforcing Sparse Video Reasoning from Dense References', Available at: https://omanscience.com/ar/articles/less-from-more-reinforcing-sparse-video-reasoning-from-dense-references.
فانكوفر
Sun W, Du Y, Snoek CGM. Less from More: Reinforcing Sparse Video Reasoning from Dense References. https://omanscience.com/ar/articles/less-from-more-reinforcing-sparse-video-reasoning-from-dense-references
IEEE
W. Sun, Y. Du, and C. G. M. Snoek, "Less from More: Reinforcing Sparse Video Reasoning from Dense References," https://omanscience.com/ar/articles/less-from-more-reinforcing-sparse-video-reasoning-from-dense-references.