الملخص
Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedural context from RGB-only demonstrations, actively select camera viewpoints, issue metric Cartesian commands, and revise them from execution feedback. Models receive no privileged object poses, oracle trajectories, or learned action heads. A fixed model-agnostic controller executes only model-specified targets. VA-Bench contains 14 base task families (11 single-arm and three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track. We evaluate 12 primary model conditions in three independent runs over the same 20 physically verified seeds per base task, reporting terminal success, nine trajectory-level behavioral diagnostics, and subtask progress. First, the best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run. Its three-run macro-average task success is only 53.93+/-3.17%. Second, active camera control significantly improves task success over passive multi-view observation. In one matched comparison, success rises from 27.86% to 57.50%. Third, held-out geometric transfer can reduce task success by over 30 percentage points. No model completes a strict long-horizon episode, despite substantial partial progress. VA-Bench thus tests whether general-purpose MLLMs can turn visual demonstrations and actively acquired evidence into successful embodied action.
الكلمات المفتاحية
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Zhang, Z., Jin, J., Wang, Y., Zhang, Z., Diao, H., Wang, L., & Lu, H. (2026). VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control. https://omanscience.com/ar/articles/vabench-measuring-embodied-spatial-intelligence-through-visual-demonstrations-active-perception-and-metric-control
MLA 9
Zhang, Zhongbo, et al. "VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control." https://omanscience.com/ar/articles/vabench-measuring-embodied-spatial-intelligence-through-visual-demonstrations-active-perception-and-metric-control.
شيكاغو (المؤلف–التاريخ)
Zhang, Zhongbo, Jiayi Jin, Yifan Wang, Zaibin Zhang, Haiwen Diao, Lijun Wang, and Huchuan Lu. 2026. "VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control." https://omanscience.com/ar/articles/vabench-measuring-embodied-spatial-intelligence-through-visual-demonstrations-active-perception-and-metric-control.
هارفارد
Zhang, Z., Jin, J., Wang, Y., Zhang, Z., Diao, H., Wang, L. and Lu, H. (2026) 'VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control', Available at: https://omanscience.com/ar/articles/vabench-measuring-embodied-spatial-intelligence-through-visual-demonstrations-active-perception-and-metric-control.
فانكوفر
Zhang Z, Jin J, Wang Y, Zhang Z, Diao H, Wang L, et al. VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control. https://omanscience.com/ar/articles/vabench-measuring-embodied-spatial-intelligence-through-visual-demonstrations-active-perception-and-metric-control
IEEE
Z. Zhang, J. Jin, Y. Wang, Z. Zhang, H. Diao, L. Wang, and H. Lu, "VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control," https://omanscience.com/ar/articles/vabench-measuring-embodied-spatial-intelligence-through-visual-demonstrations-active-perception-and-metric-control.