الملخص
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Han, Q., Hu, K., Qiu, L., Wu, C., & He, K. (2026). VISTA: A Visual Harness for Reasoning in an Interactive World. https://omanscience.com/ar/articles/vista-a-visual-harness-for-reasoning-in-an-interactive-world
MLA 9
Han, Qiushi, et al. "VISTA: A Visual Harness for Reasoning in an Interactive World." https://omanscience.com/ar/articles/vista-a-visual-harness-for-reasoning-in-an-interactive-world.
شيكاغو (المؤلف–التاريخ)
Han, Qiushi, Keya Hu, Linlu Qiu, Cathy Wu, and Kaiming He. 2026. "VISTA: A Visual Harness for Reasoning in an Interactive World." https://omanscience.com/ar/articles/vista-a-visual-harness-for-reasoning-in-an-interactive-world.
هارفارد
Han, Q., Hu, K., Qiu, L., Wu, C. and He, K. (2026) 'VISTA: A Visual Harness for Reasoning in an Interactive World', Available at: https://omanscience.com/ar/articles/vista-a-visual-harness-for-reasoning-in-an-interactive-world.
فانكوفر
Han Q, Hu K, Qiu L, Wu C, He K. VISTA: A Visual Harness for Reasoning in an Interactive World. https://omanscience.com/ar/articles/vista-a-visual-harness-for-reasoning-in-an-interactive-world
IEEE
Q. Han, K. Hu, L. Qiu, C. Wu, and K. He, "VISTA: A Visual Harness for Reasoning in an Interactive World," https://omanscience.com/ar/articles/vista-a-visual-harness-for-reasoning-in-an-interactive-world.