الباحثون

Anirudh Goyal

المنشورات 5

نسخة أولية وصول مفتوح

Agent Plasticity: Measuring Self-Improvement Through Experience

Harman Singh, Anton Bakhtin, Rulin Shao وآخرون · 2026

AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future perfor …

نسخة أولية وصول مفتوح

GitSwarm: Decentralized Compounding Inference

Vedant Shah, Ankur Samanta, Paras Dahal وآخرون · 2026

Long-horizon problem solving and scientific research require computation to accumulate across successive attempts. Partial solutions, experimental findings, and unsuccessful approaches can inform later work, yet most inference-time computation is organized around individual trajectories or candidates rather than a pers …

نسخة أولية وصول مفتوح

Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents

Ankur Samanta, Yonathan Efroni, Paul Sajda وآخرون · 2026

Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emerge several investigations later. We introduce Meta-reasoning for Iterative Research Agen …

نسخة أولية وصول مفتوح

Scaling Laws for Looped Mixture of Experts

Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introdu …

نسخة أولية وصول مفتوح

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Paras Dahal, Anton Bakhtin, Taco Cohen وآخرون · 2026

As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these ch …

المؤلفون المشاركون