الباحثون

Yingjie He

المنشورات 1

نسخة أولية وصول مفتوح

Behavior Pack Optimization for Video MLLM Post-Training

Zhaolu Kang, Shiyu Liu, Tailong Luo وآخرون · 2026

Video multimodal large language models (MLLMs) keep climbing video question answering benchmarks, yet shuffling the frames, masking the segment that supports the answer, or occluding the target object barely changes their predictions. The accuracy rests on appearance and language priors, not on the temporal evidence th …

المؤلفون المشاركون