الباحثون

Dianhai Yu

المنشورات 2

نسخة أولية وصول مفتوح

RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling

Can Zhang, Xiaotian Han, Junyuan Shang وآخرون · 2026

Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals that encoding produced, but in their deployed form each predictive frame is still tokenize …

نسخة أولية وصول مفتوح

CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

Yulong Liu, Xiaotian Han, Junyuan Shang وآخرون · 2026

Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mism …

المؤلفون المشاركون