الباحثون

Bolin Ni

المنشورات 1

نسخة أولية وصول مفتوح

How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

Lin Chen, Bolin Ni, Qi Yang وآخرون · 2026

Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterize …

المؤلفون المشاركون