الباحثون

Koushik Sen

المنشورات 1

نسخة أولية وصول مفتوح

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

Shuangjie Yao, Hao Wang, Koushik Sen وآخرون · 2026

Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks. Yet nearly all existing benchmarks still rely on the same decades-old criterion: a solution is correct if it passes a fixed set of unit tests. Such tests are often insufficient: they check o …

المؤلفون المشاركون