Abstract

Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mostly limited to snippets or functions. We introduce SWE-Flux, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs. The benchmark covers singletest and multi-test questions over control flow, loops, program state, dataflow, exceptions, and program invariants. Evaluating five LLMs shows that this task remains challenging. The best model achieves only 37% accuracy. Models perform better on localized behavior such as invariants, intra-procedural control flow, exceptions, and simple loops, but struggle with dataflow, inter-procedural execution, precise state reasoning, and suite-level aggregation. Finally, we show that the oracle-harvesting pipeline can generate fresh benchmark variants using input perturbation. It successfully harvests valid variants for almost 90% of the selected instances, and the resulting variants are substantially more challenging for the evaluated models.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Taherkhani, H., Abdollahi, M., Sepidband, M., Dhulipala, H., Nguyen, T. N., & Hemmati, H. (2026). Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark. https://omanscience.com/en/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark

MLA 9

Taherkhani, Hamed, et al. "Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark." https://omanscience.com/en/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark.

Chicago (author–date)

Taherkhani, Hamed, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, and Hadi Hemmati. 2026. "Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark." https://omanscience.com/en/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark.

Harvard

Taherkhani, H., Abdollahi, M., Sepidband, M., Dhulipala, H., Nguyen, T. N. and Hemmati, H. (2026) 'Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark', Available at: https://omanscience.com/en/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark.

Vancouver

Taherkhani H, Abdollahi M, Sepidband M, Dhulipala H, Nguyen TN, Hemmati H. Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark. https://omanscience.com/en/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark

IEEE

H. Taherkhani, M. Abdollahi, M. Sepidband, H. Dhulipala, T. N. Nguyen, and H. Hemmati, "Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark," https://omanscience.com/en/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark.