الملخص

Vision-language models (VLMs) are increasingly evaluated for egocentric and cross-view video reasoning, yet existing benchmarks largely focus on semantic event understanding, temporal relations, or correspondence between already observed views, leaving their ability to reason directly about future visual states underexplored. We introduce EgoExo-Next, a visual-option benchmark for dynamic visual-state reasoning, where models must identify how an observed action trajectory subsequently appears rather than predict only an action label or textual description. EgoExo-Next contains 2,503 human-curated four-choice questions from six public egocentric and ego--exo video sources and comprises four interconnected subtasks that evaluate egocentric next-state prediction, bidirectional ego--exo state correspondence, exocentric next-state prediction, and their composition in Ego-to-Exo Next-State. Extensive evaluation of proprietary, open-source, and spatial reasoning VLMs reveals a substantial human--model gap, with the best model achieving 43.81\% average accuracy compared with 98.55\% for humans, and the largest degradation occurring on the composed Ego-to-Exo task. These results suggest that current VLMs remain substantially limited in dynamic visual-state reasoning, particularly when temporal progression and cross-view reasoning must be composed. The benchmark is publicly available at \url{https://huggingface.co/datasets/yutongli2024/EgoExo-Next}.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Li, Y., Wang, M., Li, X., Fang, Y., Dong, D., & Ye, Z. (2026). EgoExo-Next:Benchmarking Vision-Language Models on Visual-Option Next-State and Cross-View Reasoning. https://omanscience.com/ar/articles/egoexo-next-benchmarking-vision-language-models-on-visual-option-next-state-and-cross-view-reasoning

MLA 9

Li, Yutong, et al. "EgoExo-Next:Benchmarking Vision-Language Models on Visual-Option Next-State and Cross-View Reasoning." https://omanscience.com/ar/articles/egoexo-next-benchmarking-vision-language-models-on-visual-option-next-state-and-cross-view-reasoning.

شيكاغو (المؤلف–التاريخ)

Li, Yutong, Molin Wang, Xiaotong Li, Yanyan Fang, Daoguo Dong, and Ziyi Ye. 2026. "EgoExo-Next:Benchmarking Vision-Language Models on Visual-Option Next-State and Cross-View Reasoning." https://omanscience.com/ar/articles/egoexo-next-benchmarking-vision-language-models-on-visual-option-next-state-and-cross-view-reasoning.

هارفارد

Li, Y., Wang, M., Li, X., Fang, Y., Dong, D. and Ye, Z. (2026) 'EgoExo-Next:Benchmarking Vision-Language Models on Visual-Option Next-State and Cross-View Reasoning', Available at: https://omanscience.com/ar/articles/egoexo-next-benchmarking-vision-language-models-on-visual-option-next-state-and-cross-view-reasoning.

فانكوفر

Li Y, Wang M, Li X, Fang Y, Dong D, Ye Z. EgoExo-Next:Benchmarking Vision-Language Models on Visual-Option Next-State and Cross-View Reasoning. https://omanscience.com/ar/articles/egoexo-next-benchmarking-vision-language-models-on-visual-option-next-state-and-cross-view-reasoning

IEEE

Y. Li, M. Wang, X. Li, Y. Fang, D. Dong, and Z. Ye, "EgoExo-Next:Benchmarking Vision-Language Models on Visual-Option Next-State and Cross-View Reasoning," https://omanscience.com/ar/articles/egoexo-next-benchmarking-vision-language-models-on-visual-option-next-state-and-cross-view-reasoning.