الباحثون

Wei Wen

المنشورات 2

نسخة أولية وصول مفتوح

EVIE: Evidence-Vector-Informed Embeddings for Visual Document Retrieval

Zifei Wang, Wei Wen, Qiang Ji وآخرون · 2026

Accurate and scalable visual document retrieval (VDR) requires both fine-grained page understanding and efficient indexing, yet existing approaches struggle to achieve both. OCR-based text retrieval adds preprocessing latency and can lose visual and structural cues needed to understand complex pages. Single-vector visi …

نسخة أولية وصول مفتوح

Mid-Training Language Models on Raw Video

Jaedong Hwang, Xiaoqian Shen, Ernie Chang وآخرون · 2026

Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid-training data for a pretrained language model. Frames are encoded in …

المؤلفون المشاركون