الباحثون

Molin Wang

المنشورات 1

نسخة أولية وصول مفتوح

EgoExo-Next:Benchmarking Vision-Language Models on Visual-Option Next-State and Cross-View Reasoning

Yutong Li, Molin Wang, Xiaotong Li وآخرون · 2026

Vision-language models (VLMs) are increasingly evaluated for egocentric and cross-view video reasoning, yet existing benchmarks largely focus on semantic event understanding, temporal relations, or correspondence between already observed views, leaving their ability to reason directly about future visual states underex …

المؤلفون المشاركون