الملخص
Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for evaluating interactive visual intelligence across more than 5K open-source video games curated from PyWeek and itch.io. Spanning diverse genres and engines, including Pygame, HTML5, Godot, and Unity, these independent games are largely out-of-distribution for current models, reducing the likelihood that success can be achieved by retrieving memorized walkthroughs or web-scale training artifacts. To enable scalable evaluation across heterogeneous titles, we develop a unified closed-loop interaction framework optimized for HPC clusters alongside a Video-LLM-as-a-judge protocol that maps observable gameplay milestones to standardized progress levels. We evaluate fourteen recent open models spanning vision-language models, computer-use agents, and vision-language-action models. Our results yield strong evidence of a perception-action gap: despite strong reasoning capabilities, current models struggle to make sustained progress and exhibit recurring failures in spatial grounding, action execution, and self-correction. PlaySuite provides a reproducible and extensible testbed for measuring progress from visual perception to goal-directed interaction, and a foundation for developing models that can act, adapt, and generalize in dynamic visual environments.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Varghese, D., Vettoruzzo, A., Simoncini, W., Callejas, M. L. A., Derakhshani, M. M., Meding, K., Vanschoren, J., & Snoek, C. G. M. (2026). PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence. https://omanscience.com/ar/articles/playsuite-a-large-scale-benchmark-for-interactive-visual-intelligence
MLA 9
Varghese, Dheeraj, et al. "PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence." https://omanscience.com/ar/articles/playsuite-a-large-scale-benchmark-for-interactive-visual-intelligence.
شيكاغو (المؤلف–التاريخ)
Varghese, Dheeraj, Anna Vettoruzzo, Walter Simoncini, Michelle Lorena Acevedo Callejas, Mohammad Mahdi Derakhshani, Kristof Meding, Joaquin Vanschoren, and Cees G. M. Snoek. 2026. "PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence." https://omanscience.com/ar/articles/playsuite-a-large-scale-benchmark-for-interactive-visual-intelligence.
هارفارد
Varghese, D., Vettoruzzo, A., Simoncini, W., Callejas, M. L. A., Derakhshani, M. M., Meding, K., Vanschoren, J. and Snoek, C. G. M. (2026) 'PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence', Available at: https://omanscience.com/ar/articles/playsuite-a-large-scale-benchmark-for-interactive-visual-intelligence.
فانكوفر
Varghese D, Vettoruzzo A, Simoncini W, Callejas MLA, Derakhshani MM, Meding K, et al. PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence. https://omanscience.com/ar/articles/playsuite-a-large-scale-benchmark-for-interactive-visual-intelligence
IEEE
D. Varghese, A. Vettoruzzo, W. Simoncini, M. L. A. Callejas, M. M. Derakhshani, K. Meding, J. Vanschoren, and C. G. M. Snoek, "PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence," https://omanscience.com/ar/articles/playsuite-a-large-scale-benchmark-for-interactive-visual-intelligence.