الباحثون

Yaoteng Tan

المنشورات 1

نسخة أولية وصول مفتوح

CheatBench: Measuring Reward Gaming in AI Agents

Long Phan, Stephen K. Yang, Jason J. Lim وآخرون · 2026

Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring s …

المؤلفون المشاركون