Preprint Open access
JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments
In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object, time, or view that s …