Preprint Open access
VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding
Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence …