الباحثون

Baishakhi Ray

المنشورات 2

نسخة أولية وصول مفتوح

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

Shuangjie Yao, Hao Wang, Koushik Sen وآخرون · 2026

Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks. Yet nearly all existing benchmarks still rely on the same decades-old criterion: a solution is correct if it passes a fixed set of unit tests. Such tests are often insufficient: they check o …

نسخة أولية وصول مفتوح

Cost-Efficient Theorem Proving via Agent Orchestration in Program Verification

Program verification establishes software correctness through machine-checkable proofs constructed in theorem provers. It's a guarantee especially valuable for code generated by large language models (LLMs), which is fluent but carries no assurance of correctness. Almost all existing provers, however, pursue pass rates …

المؤلفون المشاركون