نسخة أولية وصول مفتوح
Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents
Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emerge several investigations later. We introduce Meta-reasoning for Iterative Research Agen …