Abstract

Static QA and code-generation benchmarks only partially capture the role that large language models (LLMs) now play as coding agents and research tools. We introduce OptiArena, a budget-controlled testbed for studying whether LLMs can improve executable game-playing algorithms through five rounds of code edits within a fixed minimal scaffold and under bounded evaluator feedback and fixed resource budgets. The testbed uses two optimization regimes, surface obfuscation controls, calibrated references, held-out/stress splits, and diagnostics for degradation and exceptional failures, with LLM API cost reported separately from local evaluator wall-clock. The empirical study asks three questions: whether models can close the calibrated gap between a designated weak starter and an editable competent baseline, whether they can refine editable competent baselines without damaging them, and whether gains survive surface obfuscation controls. Across twelve frontier LLMs and five games, models improve designated weak starters more consistently than they refine editable competent baselines, with substantial variation across games and models. OptiArena provides a practical testbed for measuring bounded-resource algorithm optimization within the five-edit, fixed-scaffold setting studied here. Code is available at https://github.com/WJ-Peng/OptiArena.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Peng, W., & Wang, X. (2026). OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets? https://omanscience.com/en/articles/optiarena-can-llms-improve-executable-algorithms-under-fixed-resource-budgets

MLA 9

Peng, Wenjun, and Xinyu Wang. "OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?" https://omanscience.com/en/articles/optiarena-can-llms-improve-executable-algorithms-under-fixed-resource-budgets.

Chicago (author–date)

Peng, Wenjun, and Xinyu Wang. 2026. "OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?" https://omanscience.com/en/articles/optiarena-can-llms-improve-executable-algorithms-under-fixed-resource-budgets.

Harvard

Peng, W. and Wang, X. (2026) 'OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?', Available at: https://omanscience.com/en/articles/optiarena-can-llms-improve-executable-algorithms-under-fixed-resource-budgets.

Vancouver

Peng W, Wang X. OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets? https://omanscience.com/en/articles/optiarena-can-llms-improve-executable-algorithms-under-fixed-resource-budgets

IEEE

W. Peng, and X. Wang, "OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?," https://omanscience.com/en/articles/optiarena-can-llms-improve-executable-algorithms-under-fixed-resource-budgets.