Abstract

Coding agents now find real vulnerabilities in production software. However, bug discovery results do not measure whether agents can construct exploit primitives. We introduce KEX-bench, a benchmark for evaluating coding agents on exploit primitive generation against real operating-system kernels. KEX-bench contains 45 task instances across 40 Linux and Windows CVEs, covering kernel address leak, instruction-pointer control, heap read, heap write, and arbitrary address write. Each task runs in an isolated virtual machine, exposes controlled tools, and uses a deterministic verifier to check primitive-specific success. We evaluate state-of-the-art coding agents paired with frontier and open-weight models under fixed tool-call budgets. Without a reference proof of concept (PoC), the strongest configuration solves 1 of 20 Windows tasks (5.0%) and 14 of 25 Linux tasks (56.0%). With a reference PoC, the strongest configuration solves 31 of 45 tasks (68.9%). This highlights the gap where agents reach kernel crashes but fail to shape kernel state into exploit primitives. We release KEX-bench for reproducible research on AI-assisted exploitation at https://kex-bench.github.io.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Jang, J., Lee, G., Lee, H., Kim, K., Kim, J., Jung, J., & Zhang, L. (2026). Evaluating Coding Agents on Kernel Exploit Generation. https://omanscience.com/en/articles/evaluating-coding-agents-on-kernel-exploit-generation

MLA 9

Jang, Junyoung, et al. "Evaluating Coding Agents on Kernel Exploit Generation." https://omanscience.com/en/articles/evaluating-coding-agents-on-kernel-exploit-generation.

Chicago (author–date)

Jang, Junyoung, Gwanhyun Lee, Hwiwon Lee, Kyuheon Kim, Jongseong Kim, Jinho Jung, and Lingming Zhang. 2026. "Evaluating Coding Agents on Kernel Exploit Generation." https://omanscience.com/en/articles/evaluating-coding-agents-on-kernel-exploit-generation.

Harvard

Jang, J., Lee, G., Lee, H., Kim, K., Kim, J., Jung, J. and Zhang, L. (2026) 'Evaluating Coding Agents on Kernel Exploit Generation', Available at: https://omanscience.com/en/articles/evaluating-coding-agents-on-kernel-exploit-generation.

Vancouver

Jang J, Lee G, Lee H, Kim K, Kim J, Jung J, et al. Evaluating Coding Agents on Kernel Exploit Generation. https://omanscience.com/en/articles/evaluating-coding-agents-on-kernel-exploit-generation

IEEE

J. Jang, G. Lee, H. Lee, K. Kim, J. Kim, J. Jung, and L. Zhang, "Evaluating Coding Agents on Kernel Exploit Generation," https://omanscience.com/en/articles/evaluating-coding-agents-on-kernel-exploit-generation.