الباحثون

Zheyuan Liu

المنشورات 3

نسخة أولية وصول مفتوح

Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

Mingyue Cui, Zheyuan Liu, Yihan Zhu وآخرون · 2026

Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent pol …

نسخة أولية وصول مفتوح

CheatBench: Measuring Reward Gaming in AI Agents

Long Phan, Stephen K. Yang, Jason J. Lim وآخرون · 2026

Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring s …

نسخة أولية وصول مفتوح

Reward Hacking Challenges Oversight of Autonomous Research Agents

Yue Huang, Zhangchen Xu, Yuchen Ma وآخرون · 2026

Autonomous research agents can design experiments, evaluate results, and write reports, giving them control over both a scientific result and the evidence used to support it. This creates a risk of reward hacking: meeting the reward criteria without achieving the intended goal. We study (1) how often models reward-hack …

المؤلفون المشاركون