نسخة أولية وصول مفتوح
Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and …