الملخص
Large language model (LLM) agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a container, or an executable binary. This setup bypasses the critical environment reconstruction step, leaving a fundamental question for real-world vulnerability analysis: can an agent autonomously reconstruct the required execution environment and reproduce a vulnerability entirely from scratch? To address this gap, we present ReproBench, an evidence-grounded benchmark designed to evaluate agent capabilities in end-to-end vulnerability reproduction starting from solely a CVE identifier. ReproBench decomposes the full reproduction workflow into six distinct phases, and assesses performance on each phase independently using verifiable experimental artifacts: downloaded firmware images, unpacked binaries, granular analysis logs, and validated crash samples, among others. We instantiate ReproBench with 30 real-world IoT firmware vulnerabilities, which serve as ideal test cases for our from-scratch evaluation setting. Our evaluation demonstrates that 45.3% of test runs resort to vulnerability simulation - a prevalent remediation workaround adopted across all evaluated LLM agents - while only 5.3% of CVE-model pairs achieve successful reproduction of real-world vulnerabilities. Despite the low overall success rate, these non-trivial successful cases confirm that state-of-the-art LLM agents already possess the capacity for fully autonomous end-to-end vulnerability reproduction. Concurrently, our in-depth analysis of failed cases identifies core bottlenecks impeding LLM agents throughout the reproduction pipeline, offering actionable insights for subsequent research.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
He, L., Wu, S., Hao, H., Zhao, H., Yan, J., & Su, P. (2026). ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch. https://omanscience.com/ar/articles/reprobench-benchmarking-llm-agents-on-reproducing-vulnerability-from-scratch
MLA 9
He, Liang, et al. "ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch." https://omanscience.com/ar/articles/reprobench-benchmarking-llm-agents-on-reproducing-vulnerability-from-scratch.
شيكاغو (المؤلف–التاريخ)
He, Liang, Sheng Wu, Haomiao Hao, Hongduo Zhao, Jia Yan, and Purui Su. 2026. "ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch." https://omanscience.com/ar/articles/reprobench-benchmarking-llm-agents-on-reproducing-vulnerability-from-scratch.
هارفارد
He, L., Wu, S., Hao, H., Zhao, H., Yan, J. and Su, P. (2026) 'ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch', Available at: https://omanscience.com/ar/articles/reprobench-benchmarking-llm-agents-on-reproducing-vulnerability-from-scratch.
فانكوفر
He L, Wu S, Hao H, Zhao H, Yan J, Su P. ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch. https://omanscience.com/ar/articles/reprobench-benchmarking-llm-agents-on-reproducing-vulnerability-from-scratch
IEEE
L. He, S. Wu, H. Hao, H. Zhao, J. Yan, and P. Su, "ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch," https://omanscience.com/ar/articles/reprobench-benchmarking-llm-agents-on-reproducing-vulnerability-from-scratch.