الملخص
A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow. Manually verifying whether a proposed no-code fix resolves the reported bug takes considerable developer time. This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment. We evaluate 322 no-code fixes generated by the 12 configurations released with the benchmark of a previous study, covering bug reports categorized as Faulty Configuration, Wrong Version, or External System & Dependency. An executor agent applies each fix by following its natural-language instructions, and an issue-specific checker determines whether the reported bug persists. We repeat the pipeline with three executors: two Computer-Use Agents, OpenCUA-72B and Claude Sonnet 5, and one multimodal agentic LLM, Meta's Muse Glimmer. Only 17.6% of the candidate issues could be set up and passed both sanity gates. Across the 322 fixes, 14.6% to 49.7% resolved the bug depending on the executor, and the strongest configuration, Claude Opus 4.6 in the Vanilla pipeline, resolved up to 74.1% of its fixes under Claude Sonnet 5. Changing only the executor shifted a configuration's resolution rate by 38.8% on average, and the three executors reached the same verdict on only 46.9% of the fixes. Compared with human execution, the executors matched the human consensus for 66.1% to 88.1% of the sampled fixes. Even under the best executor, fewer than half of the LLM-generated no-code fixes resolve the reported bug, so such fixes need verification before they reach users. Execution-based verification can provide this, but the measured capability depends strongly on the executor, which evaluations must report and control.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Torun, U. B., Karakaya, V., & Tüzün, E. (2026). Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes. https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes
MLA 9
Torun, Utku Boran, et al. "Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes." https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes.
شيكاغو (المؤلف–التاريخ)
Torun, Utku Boran, Veli Karakaya, and Eray Tüzün. 2026. "Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes." https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes.
هارفارد
Torun, U. B., Karakaya, V. and Tüzün, E. (2026) 'Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes', Available at: https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes.
فانكوفر
Torun UB, Karakaya V, Tüzün E. Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes. https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes
IEEE
U. B. Torun, V. Karakaya, and E. Tüzün, "Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes," https://omanscience.com/ar/articles/can-llms-fix-it-without-code-toward-automated-verification-of-no-code-bug-fixes.