Preprint Open access
AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering
Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing these failure causes is essential to diagnose and improve agen …