الباحثون

Param Biyani

المنشورات 1

نسخة أولية وصول مفتوح

SpecGuard: Proving a Task Is Broken Before the Agent Cheats

As autonomous coding agents get increasingly deployed, the risk that accidental or adversarially injected misspecifications in tasks lead to dangerous agent behavior is critical to address. Prior work has shown that agents given such tasks rarely flag the conflict and instead cheat, editing tests or hard-coding expecte …

المؤلفون المشاركون