Abstract
Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong hypotheses, and can lead to sparse terminal reward and unstable training. We introduce Contextual Early Reward (CER), which predicts terminal reward through behavioral evidence in a trajectory prefix. CER synthesizes adaptive rubrics specific to the current task and stage through experiences summarized from related historical tasks. In test-time scaling on SWE-bench Verified, CER improves RM@8 over the strongest baseline by 4.2 percentage points (pp) on Nemotron 3 Ultra and 2.0 pp on Qwen 3.6 27B; on Nemotron, it takes only 15.3% tokens to match the best baseline performance. In RL training experiments, CER exceeds full-rollout TMax by 1.9 pp while using 52.7% fewer online policy-and-judge tokens. Together, CER provides an interpretable, efficient, and dense evaluation method for long-horizon coding agents.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Yao, J., Zeng, S., Feng, S., Fan, Z., Zhu, B., & Tsvetkov, Y. (2026). Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents. https://omanscience.com/en/articles/before-the-rollout-ends-early-terminal-reward-prediction-for-long-horizon-coding-agents
MLA 9
Yao, Jihan, et al. "Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents." https://omanscience.com/en/articles/before-the-rollout-ends-early-terminal-reward-prediction-for-long-horizon-coding-agents.
Chicago (author–date)
Yao, Jihan, Sihan Zeng, Shangbin Feng, Zhiyuan Fan, Banghua Zhu, and Yulia Tsvetkov. 2026. "Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents." https://omanscience.com/en/articles/before-the-rollout-ends-early-terminal-reward-prediction-for-long-horizon-coding-agents.
Harvard
Yao, J., Zeng, S., Feng, S., Fan, Z., Zhu, B. and Tsvetkov, Y. (2026) 'Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents', Available at: https://omanscience.com/en/articles/before-the-rollout-ends-early-terminal-reward-prediction-for-long-horizon-coding-agents.
Vancouver
Yao J, Zeng S, Feng S, Fan Z, Zhu B, Tsvetkov Y. Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents. https://omanscience.com/en/articles/before-the-rollout-ends-early-terminal-reward-prediction-for-long-horizon-coding-agents
IEEE
J. Yao, S. Zeng, S. Feng, Z. Fan, B. Zhu, and Y. Tsvetkov, "Before the Rollout Ends: Early Terminal Reward Prediction for Long-horizon Coding Agents," https://omanscience.com/en/articles/before-the-rollout-ends-early-terminal-reward-prediction-for-long-horizon-coding-agents.