Abstract

Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a codec-native vision encoder with native token compression for unified image-video understanding. By combining codec-native input support with segmented attention, CoVisco can encode long visual inputs in a single forward pass without forming dense patch-to-patch interactions across all frames. Each temporal segment is equipped with learnable abstract tokens that learn a compact segment-level representation, while fine-grained patch tokens remain available throughout the encoder. Alternating intra-segment and abstract-communication layers preserve video-level context through the abstract-token channel. A lightweight selector further exposes either abstract tokens alone or abstract tokens augmented with a runtime-selected subset of patch tokens, yielding a compact visual interface that reduces the visual context and prefill burden of downstream MLLMs while retaining fine-grained evidence when needed. Pretrained with contrastive objectives on 565M image--text pairs and 6.4M videos, CoVisco shows competitive performance on video-oriented embedding and multimodal understanding benchmarks. In the evaluated four-segment, 64-frame setting, abstract-only inference uses only 400 visual tokens while achieving video-understanding performance close to, and on some benchmarks exceeding, OneVision-Encoder. Selected patch tokens further improve fine-grained video reasoning. Project URL: https://github.com/ernie-research/CoVisco.git

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Liu, Y., Han, X., Shang, J., Ding, Y., Zhang, Z., Wang, S., Zhu, G., Han, S., & Yu, D. (2026). CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding. https://omanscience.com/en/articles/covisco-codec-native-vision-encoder-with-native-token-compression-for-unified-image-video-understanding

MLA 9

Liu, Yulong, et al. "CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding." https://omanscience.com/en/articles/covisco-codec-native-vision-encoder-with-native-token-compression-for-unified-image-video-understanding.

Chicago (author–date)

Liu, Yulong, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Guibo Zhu, Sirui Han, and Dianhai Yu. 2026. "CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding." https://omanscience.com/en/articles/covisco-codec-native-vision-encoder-with-native-token-compression-for-unified-image-video-understanding.

Harvard

Liu, Y., Han, X., Shang, J., Ding, Y., Zhang, Z., Wang, S., Zhu, G., Han, S. and Yu, D. (2026) 'CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding', Available at: https://omanscience.com/en/articles/covisco-codec-native-vision-encoder-with-native-token-compression-for-unified-image-video-understanding.

Vancouver

Liu Y, Han X, Shang J, Ding Y, Zhang Z, Wang S, et al. CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding. https://omanscience.com/en/articles/covisco-codec-native-vision-encoder-with-native-token-compression-for-unified-image-video-understanding

IEEE

Y. Liu, X. Han, J. Shang, Y. Ding, Z. Zhang, S. Wang, G. Zhu, S. Han, and D. Yu, "CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding," https://omanscience.com/en/articles/covisco-codec-native-vision-encoder-with-native-token-compression-for-unified-image-video-understanding.