الملخص
Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mostly limited to snippets or functions. We introduce SWE-Flux, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs. The benchmark covers singletest and multi-test questions over control flow, loops, program state, dataflow, exceptions, and program invariants. Evaluating five LLMs shows that this task remains challenging. The best model achieves only 37% accuracy. Models perform better on localized behavior such as invariants, intra-procedural control flow, exceptions, and simple loops, but struggle with dataflow, inter-procedural execution, precise state reasoning, and suite-level aggregation. Finally, we show that the oracle-harvesting pipeline can generate fresh benchmark variants using input perturbation. It successfully harvests valid variants for almost 90% of the selected instances, and the resulting variants are substantially more challenging for the evaluated models.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Taherkhani, H., Abdollahi, M., Sepidband, M., Dhulipala, H., Nguyen, T. N., & Hemmati, H. (2026). Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark. https://omanscience.com/ar/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark
MLA 9
Taherkhani, Hamed, et al. "Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark." https://omanscience.com/ar/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark.
شيكاغو (المؤلف–التاريخ)
Taherkhani, Hamed, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, and Hadi Hemmati. 2026. "Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark." https://omanscience.com/ar/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark.
هارفارد
Taherkhani, H., Abdollahi, M., Sepidband, M., Dhulipala, H., Nguyen, T. N. and Hemmati, H. (2026) 'Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark', Available at: https://omanscience.com/ar/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark.
فانكوفر
Taherkhani H, Abdollahi M, Sepidband M, Dhulipala H, Nguyen TN, Hemmati H. Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark. https://omanscience.com/ar/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark
IEEE
H. Taherkhani, M. Abdollahi, M. Sepidband, H. Dhulipala, T. N. Nguyen, and H. Hemmati, "Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark," https://omanscience.com/ar/articles/can-llms-reason-about-runtime-behavior-a-repository-level-dynamic-benchmark.