نسخة أولية وصول مفتوح
Cross-Benchmark Transfer from RL on Agentic Coding Tasks
Coding agents often fail in the last mile: they build most of a feature but drop a requirement, test only the cases their implementation already handles, break behavior that was supposed to stay intact, or validate against an unchecked assumption. We ask whether reinforcement learning (RL) on expert-built agentic codin …