الباحثون

Huanfen Yao

المنشورات 2

نسخة أولية وصول مفتوح

Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis

Xia Hu, Brian Potetz, Chun-Ta Lu وآخرون · 2026

End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through control …

نسخة أولية وصول مفتوح

Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows

Xiyuan Shen, Jiuyang Lyu, Seokhyun Hwang وآخرون · 2026

Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language models (VLMs) become increasingly capable of analyzing video, they offer opportunities to automate this traditionally human- …

المؤلفون المشاركون