الملخص

A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow. Manually verifying whether a proposed no-code fix resolves the reported bug takes considerable developer time. This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment. We evaluate 322 no-code fixes generated by the 12 configurations released with the benchmark of a previous study, covering bug reports categorized as Faulty Configuration, Wrong Version, or External System & Dependency. An executor agent applies each fix by following its natural-language instructions, and an issue-specific checker determines whether the reported bug persists. We repeat the pipeline with three executors: two Computer-Use Agents, OpenCUA-72B and Claude Sonnet 5, and one multimodal agentic LLM, Meta's Muse Glimmer. Only 17.6% of the candidate issues could be set up and passed both sanity gates. Across the 322 fixes, 14.6% to 49.7% resolved the bug depending on the executor, and the strongest configuration, Claude Opus 4.6 in the Vanilla pipeline, resolved up to 74.1% of its fixes under Claude Sonnet 5. Changing only the executor shifted a configuration's resolution rate by 38.8% on average, and the three executors reached the same verdict on only 46.9% of the fixes. Compared with human execution, the executors matched the human consensus for 66.1% to 88.1% of the sampled fixes. Even under the best executor, fewer than half of the LLM-generated no-code fixes resolve the reported bug, so such fixes need verification before they reach users. Execution-based verification can provide this, but the measured capability depends strongly on the executor, which evaluations must report and control.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Torun, U. B., Karakaya, V., & Tüzün, E. (2026). Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes. https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes

MLA 9

Torun, Utku Boran, et al. "Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes." https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes.

شيكاغو (المؤلف–التاريخ)

Torun, Utku Boran, Veli Karakaya, and Eray Tüzün. 2026. "Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes." https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes.

هارفارد

Torun, U. B., Karakaya, V. and Tüzün, E. (2026) 'Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes', Available at: https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes.

فانكوفر

Torun UB, Karakaya V, Tüzün E. Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes. https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes

IEEE

U. B. Torun, V. Karakaya, and E. Tüzün, "Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes," https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes.