Abstract
Autonomous research agents can design experiments, evaluate results, and write reports, giving them control over both a scientific result and the evidence used to support it. This creates a risk of reward hacking: meeting the reward criteria without achieving the intended goal. We study (1) how often models reward-hack without instructions to do so, (2) how effective and detectable their methods are when hacking is allowed, and (3) how they adapt when an LLM review panel returns its decision and reasons. Across 17 language models and 38 tasks, the spontaneous reward-hacking rate is 30.5% on open-ended research-pipeline tasks and 2.9% on task-specific kernels. When hacking is allowed on tasks whose pass thresholds exceed our best compliant baselines, 505/677 attempts (74.6%) are confirmed reward hacks: they both clear the threshold and receive mechanism-verification panel confirmation of an evaluation exploit. An LLM panel reviewing only submitted code and reported scores misses 33/505 confirmed hacks (6.5%). Direct methods that achieve the highest scores are often easy to detect, while less direct methods evade more often. In a five-round loop, the number of model-task pairs with an evasion rises from 7 to 56. Among 79 pairs evaluated under two feedback conditions, cumulative evasion reaches 40.5% with detailed feedback and 20.3% with generic rejection. The detailed condition includes the review decision, reasons, and attempt history, so this comparison does not isolate the effect of explanations. These findings highlight the need for stronger defenses, including metrics kept outside the agent's control and independent recomputation on data chosen to expose likely exploits.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Huang, Y., Xu, Z., Ma, Y., Wang, W., Liu, Z., Xu, Z., Chen, P. Y., Galley, M., Lin, Z., Feuerriegel, S., Poovendran, R., Sra, M., Pentland, A., Zhang, X., & Chen, Z. (2026). Reward Hacking Challenges Oversight of Autonomous Research Agents. https://omanscience.com/en/articles/reward-hacking-challenges-oversight-of-autonomous-research-agents
MLA 9
Huang, Yue, et al. "Reward Hacking Challenges Oversight of Autonomous Research Agents." https://omanscience.com/en/articles/reward-hacking-challenges-oversight-of-autonomous-research-agents.
Chicago (author–date)
Huang, Yue, Zhangchen Xu, Yuchen Ma, Wenjie Wang, Zheyuan Liu, Ziwei Xu, Pin-Yu Chen, Michel Galley, Zinan Lin, Stefan Feuerriegel, Radha Poovendran, Misha Sra, Alex Pentland, Xiangliang Zhang, and Zichen Chen. 2026. "Reward Hacking Challenges Oversight of Autonomous Research Agents." https://omanscience.com/en/articles/reward-hacking-challenges-oversight-of-autonomous-research-agents.
Harvard
Huang, Y., Xu, Z., Ma, Y., Wang, W., Liu, Z., Xu, Z., Chen, P. Y., Galley, M., Lin, Z., Feuerriegel, S., Poovendran, R., Sra, M., Pentland, A., Zhang, X. and Chen, Z. (2026) 'Reward Hacking Challenges Oversight of Autonomous Research Agents', Available at: https://omanscience.com/en/articles/reward-hacking-challenges-oversight-of-autonomous-research-agents.
Vancouver
Huang Y, Xu Z, Ma Y, Wang W, Liu Z, Xu Z, et al. Reward Hacking Challenges Oversight of Autonomous Research Agents. https://omanscience.com/en/articles/reward-hacking-challenges-oversight-of-autonomous-research-agents
IEEE
Y. Huang, Z. Xu, Y. Ma, W. Wang, Z. Liu, Z. Xu, P. Y. Chen, M. Galley, Z. Lin, S. Feuerriegel, R. Poovendran, M. Sra, A. Pentland, X. Zhang, and Z. Chen, "Reward Hacking Challenges Oversight of Autonomous Research Agents," https://omanscience.com/en/articles/reward-hacking-challenges-oversight-of-autonomous-research-agents.