الباحثون

Raghuraman Krishnamoorthi

المنشورات 2

نسخة أولية وصول مفتوح

Mid-Training Language Models on Raw Video

Jaedong Hwang, Xiaoqian Shen, Ernie Chang وآخرون · 2026

Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid-training data for a pretrained language model. Frames are encoded in …

نسخة أولية وصول مفتوح

Scaling Laws for Looped Mixture of Experts

Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introdu …

المؤلفون المشاركون