الباحثون

Sihan Zeng

المنشورات 1

نسخة أولية وصول مفتوح

Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents

Jihan Yao, Sihan Zeng, Shangbin Feng وآخرون · 2026

Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong hypotheses, and can lead to sparse terminal reward and unstable training. We introduce Contextual Early Reward (CER), which predicts terminal reward through …

المؤلفون المشاركون