Preprint Open access
RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling
Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals that encoding produced, but in their deployed form each predictive frame is still tokenize …