الملخص
Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals that encoding produced, but in their deployed form each predictive frame is still tokenized on its own: the tokens are a function of the current primitives, not of a carried reference. We argue that a more natural function is of both---the current primitives and a carried reference. A clip and its time reversal share the same frames and differ only in the order of changes---an axis that symmetric pooling discards by construction, and that is non-empty in the frozen vision features VideoLMs use---and the codec recurrence already composes those changes in order against a reference state. We introduce RESUME, a stateful codec representation: an anchor I-frame initializes a compact latent state, each subsequent predictive frame is consumed as an update to that state, and a shared readout exposes VideoLM-compatible tokens from the accumulated state. Codec prediction is thereby kept at the representation level and handed to the language model as a trajectory, not as a set of independent token groups. At the same per-predictive-frame token budget as prior codec-aware methods, a predictive frame enters the language model as a readout of what the front-end already knows, not as an encoding of the current primitives alone. Across ten benchmarks, the gains concentrate on temporal reasoning: on all three temporal benchmarks RESUME improves over both the RGB-frame baseline LLaVA-Video-7B (by 2.8, 5.1, and 3.9 points on TempCompass, TOMATO, and MVBench) and the codec-based baseline CoPE-7B, while staying competitive on general and long-form QA. Frozen-transition tests further show anchor dependence, order sensitivity, and useful rollout behavior beyond the training horizon.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Zhang, C., Han, X., Shang, J., Ding, Y., Zhang, Z., Wang, S., Yu, D., & Li, R. (2026). RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling. https://omanscience.com/ar/articles/resume-recurrent-state-updates-from-motion-and-residual-signals-for-efficient-video-language-modeling
MLA 9
Zhang, Can, et al. "RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling." https://omanscience.com/ar/articles/resume-recurrent-state-updates-from-motion-and-residual-signals-for-efficient-video-language-modeling.
شيكاغو (المؤلف–التاريخ)
Zhang, Can, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Dianhai Yu, and Ruirui Li. 2026. "RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling." https://omanscience.com/ar/articles/resume-recurrent-state-updates-from-motion-and-residual-signals-for-efficient-video-language-modeling.
هارفارد
Zhang, C., Han, X., Shang, J., Ding, Y., Zhang, Z., Wang, S., Yu, D. and Li, R. (2026) 'RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling', Available at: https://omanscience.com/ar/articles/resume-recurrent-state-updates-from-motion-and-residual-signals-for-efficient-video-language-modeling.
فانكوفر
Zhang C, Han X, Shang J, Ding Y, Zhang Z, Wang S, et al. RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling. https://omanscience.com/ar/articles/resume-recurrent-state-updates-from-motion-and-residual-signals-for-efficient-video-language-modeling
IEEE
C. Zhang, X. Han, J. Shang, Y. Ding, Z. Zhang, S. Wang, D. Yu, and R. Li, "RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling," https://omanscience.com/ar/articles/resume-recurrent-state-updates-from-motion-and-residual-signals-for-efficient-video-language-modeling.