نسخة أولية وصول مفتوح
Pressure, Context, and Machine Self-Control: A Criminological Test of Reward Hacking in Generative AI Models
Recent incidents show that AI agents sometimes reach measured goals through unsanctioned means. This study applies self-control, general strain, anomie, neutralization and routine activity theory to reward hacking in generative AI models, and it treats the measures as behavioral analogues. Study 1 (2,310 conversations, …