الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Dahal, P., Bakhtin, A., Cohen, T., Chen, Z., Wu, C. J., Fergus, R., Yih, S., Synnaeve, G., Salakhutdinov, R., Arora, S., Weston, J., & Goyal, A. (2026). التفكير قبل التفكير: توسيع الاستدلال الوكيل عبر الاستدلال الفوقي. https://omanscience.com/ar/articles/thinking-before-thinking-scaling-agentic-inference-through-meta-reasoning

MLA 9

Dahal, Paras, et al. "التفكير قبل التفكير: توسيع الاستدلال الوكيل عبر الاستدلال الفوقي." https://omanscience.com/ar/articles/thinking-before-thinking-scaling-agentic-inference-through-meta-reasoning.

شيكاغو (المؤلف–التاريخ)

Dahal, Paras, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu, Rob Fergus, Scott Yih, Gabriel Synnaeve, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, and Anirudh Goyal. 2026. "التفكير قبل التفكير: توسيع الاستدلال الوكيل عبر الاستدلال الفوقي." https://omanscience.com/ar/articles/thinking-before-thinking-scaling-agentic-inference-through-meta-reasoning.

هارفارد

Dahal, P., Bakhtin, A., Cohen, T., Chen, Z., Wu, C. J., Fergus, R., Yih, S., Synnaeve, G., Salakhutdinov, R., Arora, S., Weston, J. and Goyal, A. (2026) 'التفكير قبل التفكير: توسيع الاستدلال الوكيل عبر الاستدلال الفوقي', Available at: https://omanscience.com/ar/articles/thinking-before-thinking-scaling-agentic-inference-through-meta-reasoning.

فانكوفر

Dahal P, Bakhtin A, Cohen T, Chen Z, Wu CJ, Fergus R, et al. التفكير قبل التفكير: توسيع الاستدلال الوكيل عبر الاستدلال الفوقي. https://omanscience.com/ar/articles/thinking-before-thinking-scaling-agentic-inference-through-meta-reasoning

IEEE

P. Dahal, A. Bakhtin, T. Cohen, Z. Chen, C. J. Wu, R. Fergus, S. Yih, G. Synnaeve, R. Salakhutdinov, S. Arora, J. Weston, and A. Goyal, "التفكير قبل التفكير: توسيع الاستدلال الوكيل عبر الاستدلال الفوقي," https://omanscience.com/ar/articles/thinking-before-thinking-scaling-agentic-inference-through-meta-reasoning.