Abstract

When post-training large language models on tasks with semi-verifiable rewards, there are many factors (training steps, base model size, training order, data quality, verifier accuracy, etc.) that practitioners must contend with to maximize model performance. Yet, it remains unclear how well verifier agreement predicts post-training performance on such tasks. In this paper, we explore this question with over 11k H100 GPU-hours, across HealthBench and PRBench tasks in medical, legal, and finance domains. Across the tested domains, Qwen3 trainees (1.7B-8B on HealthBench; 8B on PRBench), evaluation splits, and frontier LLM reference judges (which we call golden verifiers), higher verifier agreement does not consistently identify the best training verifier. Expensive verifiers need not outperform inexpensive ones, and open-weight Gemma verifiers produce strong training outcomes. We compare two low-cost choices retrospectively -- a cost-reducing choice and a balanced choice -- with estimated grading cost reductions of 98.8%-99.7% relative to the golden grading protocols and average post-training score gaps of 1-3 points from the best evaluated training verifier. These averages include larger losses in individual settings; they do not establish that verifier choices are interchangeable.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Plesner, A., Northcutt, C., Guzmán, F., & Athalye, A. (2026). A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards. https://omanscience.com/en/articles/a-cheap-verifier-is-good-enough-llm-post-training-is-robust-to-erroneous-rewards

MLA 9

Plesner, Andreas, et al. "A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards." https://omanscience.com/en/articles/a-cheap-verifier-is-good-enough-llm-post-training-is-robust-to-erroneous-rewards.

Chicago (author–date)

Plesner, Andreas, Curtis Northcutt, Francisco Guzmán, and Anish Athalye. 2026. "A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards." https://omanscience.com/en/articles/a-cheap-verifier-is-good-enough-llm-post-training-is-robust-to-erroneous-rewards.

Harvard

Plesner, A., Northcutt, C., Guzmán, F. and Athalye, A. (2026) 'A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards', Available at: https://omanscience.com/en/articles/a-cheap-verifier-is-good-enough-llm-post-training-is-robust-to-erroneous-rewards.

Vancouver

Plesner A, Northcutt C, Guzmán F, Athalye A. A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards. https://omanscience.com/en/articles/a-cheap-verifier-is-good-enough-llm-post-training-is-robust-to-erroneous-rewards

IEEE

A. Plesner, C. Northcutt, F. Guzmán, and A. Athalye, "A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards," https://omanscience.com/en/articles/a-cheap-verifier-is-good-enough-llm-post-training-is-robust-to-erroneous-rewards.