نسخة أولية وصول مفتوح
MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses
Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving beyond the evolution o …