الباحثون

Qi Qian

المنشورات 2

نسخة أولية وصول مفتوح

Mid-Training Language Models on Raw Video

Jaedong Hwang, Xiaoqian Shen, Ernie Chang وآخرون · 2026

Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid-training data for a pretrained language model. Frames are encoded in …

نسخة أولية وصول مفتوح

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

Sixun Dong, Wei Li, Andong Deng وآخرون · 2026

Efficient long-video understanding with vision-language models (VLMs) is often framed as selecting informative frames or visual tokens at a fixed native resolution. We show that per-frame resolution can instead be traded for denser temporal coverage, while front-end decoding latency depends on the size of the candidate …

المؤلفون المشاركون