نسخة أولية وصول مفتوح
Less from More: Reinforcing Sparse Video Reasoning from Dense References
Video-language models commonly assume that more temporal observations lead to more reliable reasoning. We question this assumption and argue that the key challenge is not merely processing more video frames efficiently, but learning to reason reliably under limited temporal evidence. We propose SAVER, a dense-to-sparse …