نسخة أولية وصول مفتوح
Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@$K$ tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation requ …