الباحثون

Tan Yu

المنشورات 2

نسخة أولية وصول مفتوح

Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models

Tan Yu, Alexander Bukharin, Khushi Bhardwaj وآخرون · 2026

How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@$K$ tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation requ …

نسخة أولية وصول مفتوح

LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models

Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. We identify a recurring reasoning failure, stable-but-wrong lock-in, where an answer stabilizes early around an incorrect value while substantial denoising remains. Surface-level de …

المؤلفون المشاركون