Abstract
Video-text pretraining has achieved remarkable progress through the scaling of models and datasets, yet the quality of language supervision remains underexplored. Existing web-scale datasets often provide only a single sparse caption per video that fails to capture rich spatiotemporal semantics, while directly using captioning models can generate noisy descriptions. We propose a large-scale multimodal large language model-based supervision generation framework that improves supervision diversity, fidelity, and semantic coverage. Starting from 10 million videos, our approach generates multi-view captions (MVC) through complementary summary and detailed captions, reasoning-based refinement, and semantic positive caption generation. To effectively exploit supervision at different granularities, we further introduce a granularity-aware text representation with separate CLS tokens for summary and detailed views. We pretrain video-text models using the resulting supervision corpus and evaluate them across standard, fine-grained and detailed text-to-video retrieval benchmarks. Our approach consistently improves both zero-shot and fine-tuned performance while using smaller pretraining corpora than existing methods, demonstrating the importance of rich and complementary textual supervision for video-text pretraining. Project page: https://rvandeghen.github.io/mvc/
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Thoker, F. M., Vandeghen, R., Sanchez, K., Van Droogenbroeck, M., & Ghanem, B. (2026). Advancing Video-Text Pretraining with Multi-View Captions. https://omanscience.com/en/articles/advancing-video-text-pretraining-with-multi-view-captions
MLA 9
Thoker, Fida M., et al. "Advancing Video-Text Pretraining with Multi-View Captions." https://omanscience.com/en/articles/advancing-video-text-pretraining-with-multi-view-captions.
Chicago (author–date)
Thoker, Fida M., Renaud Vandeghen, Karen Sanchez, Marc Van Droogenbroeck, and Bernard Ghanem. 2026. "Advancing Video-Text Pretraining with Multi-View Captions." https://omanscience.com/en/articles/advancing-video-text-pretraining-with-multi-view-captions.
Harvard
Thoker, F. M., Vandeghen, R., Sanchez, K., Van Droogenbroeck, M. and Ghanem, B. (2026) 'Advancing Video-Text Pretraining with Multi-View Captions', Available at: https://omanscience.com/en/articles/advancing-video-text-pretraining-with-multi-view-captions.
Vancouver
Thoker FM, Vandeghen R, Sanchez K, Van Droogenbroeck M, Ghanem B. Advancing Video-Text Pretraining with Multi-View Captions. https://omanscience.com/en/articles/advancing-video-text-pretraining-with-multi-view-captions
IEEE
F. M. Thoker, R. Vandeghen, K. Sanchez, M. Van Droogenbroeck, and B. Ghanem, "Advancing Video-Text Pretraining with Multi-View Captions," https://omanscience.com/en/articles/advancing-video-text-pretraining-with-multi-view-captions.