الملخص
Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate five parts of the agent's configuration: task information, reasoning, self-verification, time budget, and backbone model. We find substantial run-to-run variability, with approximately 54% of the outcome variance coming from repeating the same configuration rather than changing it. Across configurations, the information provided to the agent has the largest effect, exceeding both time budget and model size, while also reducing cost and improving calibration. Configuration choices also interact: additional time helps only when the agent has sufficient information or a capable enough model to use it. Finally, a trajectory-based taxonomy of agent behavior reveals that prompting an agent to verify its answer has little effect on its verification behavior, whereas providing a dedicated verification tool changes that behavior substantially. These results suggest that agents should be evaluated as configurable systems themselves, and that some desired behaviors are more effectively implemented in the system than requested through prompting. We release the benchmark and more than 18,000 agent trajectories.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wiedmann, L., Girrbach, L., Schmid, C., & Akata, Z. (2026). Agents Are Systems, Not Models: Rethinking Agentic Evaluation. https://omanscience.com/ar/articles/agents-are-systems-not-models-rethinking-agentic-evaluation
MLA 9
Wiedmann, Luis, et al. "Agents Are Systems, Not Models: Rethinking Agentic Evaluation." https://omanscience.com/ar/articles/agents-are-systems-not-models-rethinking-agentic-evaluation.
شيكاغو (المؤلف–التاريخ)
Wiedmann, Luis, Leander Girrbach, Cordelia Schmid, and Zeynep Akata. 2026. "Agents Are Systems, Not Models: Rethinking Agentic Evaluation." https://omanscience.com/ar/articles/agents-are-systems-not-models-rethinking-agentic-evaluation.
هارفارد
Wiedmann, L., Girrbach, L., Schmid, C. and Akata, Z. (2026) 'Agents Are Systems, Not Models: Rethinking Agentic Evaluation', Available at: https://omanscience.com/ar/articles/agents-are-systems-not-models-rethinking-agentic-evaluation.
فانكوفر
Wiedmann L, Girrbach L, Schmid C, Akata Z. Agents Are Systems, Not Models: Rethinking Agentic Evaluation. https://omanscience.com/ar/articles/agents-are-systems-not-models-rethinking-agentic-evaluation
IEEE
L. Wiedmann, L. Girrbach, C. Schmid, and Z. Akata, "Agents Are Systems, Not Models: Rethinking Agentic Evaluation," https://omanscience.com/ar/articles/agents-are-systems-not-models-rethinking-agentic-evaluation.