الباحثون

Qinsi Wang

المنشورات 2

نسخة أولية وصول مفتوح

Mid-Training Language Models on Raw Video

Jaedong Hwang, Xiaoqian Shen, Ernie Chang وآخرون · 2026

Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid-training data for a pretrained language model. Frames are encoded in …

نسخة أولية وصول مفتوح

Joint Branch-Space Transform Coding for Diffusion Activation Quantization with Classifier-Free Guidance

Mingrun Jiang, Yuejia Liu, Zishan Shao وآخرون · 2026

Post-training quantization for diffusion models increasingly exploits timestep, feature, and layer structure. While recent work has begun incorporating CFG structure into diffusion quantization, activation quantization still operates independently across conditional and unconditional coordinates, leaving cross-activati …

المؤلفون المشاركون