Preprint Open access
How Well Do LLMs Reason with Noisy Evidence? An Active Visual Reasoning Benchmark
Real-world reasoning rarely reduces to static question answering: agents must actively gather information from tools and sensors that are often noisy and unreliable. Yet most existing active reasoning benchmarks assume that environmental feedback is trustworthy, or introduce noise without exposing an explicit, calibrat …