الملخص

In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object, time, or view that supports it. Current visual reasoning benchmarks largely evaluate passive observations and final answers, overlooking settings that require active reasoning and evidence acquisition. We introduce JRDB-AVR, a benchmark derived from existing real-world JRDB robotics data through a structured question-generation engine that turns this gap into an explicit evaluation: an embodied agentic system receives a visual reasoning question, requests bounded observations by timestamp and viewing angle, and is evaluated on both the final answer and the grounded visual evidence supporting it. The benchmark contains diverse questions over multiple real-world environments involving temporal search, viewpoint selection, and human-oriented compositional reasoning. We also introduce JRDB-AVR-Agent, a reference active reasoning agentic method that maintains an explicit observation-grounded graph-based world model and answers through solving. Experiments reveal a substantial gap between answer accuracy and evidence accuracy in current baselines, showing that current VLMs can produce unsupported correct answers and that active evidence-aware evaluation is necessary for embodied visual reasoning. Code and benchmark are available at https://github.com/ControlNet/JRDB-AVR.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Cai, Z., Ke, F., Huang, S., Banda, M. G. D. L., Stuckey, P. J., Haffari, G., & Rezatofighi, H. (2026). JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments. https://omanscience.com/ar/articles/jrdb-avr-an-active-visual-reasoning-benchmark-for-embodied-agents-in-real-world-environments

MLA 9

Cai, Zhixi, et al. "JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments." https://omanscience.com/ar/articles/jrdb-avr-an-active-visual-reasoning-benchmark-for-embodied-agents-in-real-world-environments.

شيكاغو (المؤلف–التاريخ)

Cai, Zhixi, Fucai Ke, Sukai Huang, Maria Garcia de la Banda, Peter J. Stuckey, Gholamreza Haffari, and Hamid Rezatofighi. 2026. "JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments." https://omanscience.com/ar/articles/jrdb-avr-an-active-visual-reasoning-benchmark-for-embodied-agents-in-real-world-environments.

هارفارد

Cai, Z., Ke, F., Huang, S., Banda, M. G. D. L., Stuckey, P. J., Haffari, G. and Rezatofighi, H. (2026) 'JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments', Available at: https://omanscience.com/ar/articles/jrdb-avr-an-active-visual-reasoning-benchmark-for-embodied-agents-in-real-world-environments.

فانكوفر

Cai Z, Ke F, Huang S, Banda MGDL, Stuckey PJ, Haffari G, et al. JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments. https://omanscience.com/ar/articles/jrdb-avr-an-active-visual-reasoning-benchmark-for-embodied-agents-in-real-world-environments

IEEE

Z. Cai, F. Ke, S. Huang, M. G. D. L. Banda, P. J. Stuckey, G. Haffari, and H. Rezatofighi, "JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments," https://omanscience.com/ar/articles/jrdb-avr-an-active-visual-reasoning-benchmark-for-embodied-agents-in-real-world-environments.