الباحثون

Murat Ozer

المنشورات 2

نسخة أولية وصول مفتوح

Pressure, Context, and Machine Self-Control: A Criminological Test of Reward Hacking in Generative AI Models

Recent incidents show that AI agents sometimes reach measured goals through unsanctioned means. This study applies self-control, general strain, anomie, neutralization and routine activity theory to reward hacking in generative AI models, and it treats the measures as behavioral analogues. Study 1 (2,310 conversations, …

نسخة أولية وصول مفتوح

Reward Hacking and Agent Containment Failure: A Monte Carlo Study Based on the 2026 Hugging Face Incident

The July 2026 intrusion into Hugging Face production infrastructure showed how reward hacking can become an external cybersecurity incident when a capable agent encounters weak containment boundaries. This study develops a probabilistic risk model linking five stages: reward hacking, containment escape, usable access, …

المؤلفون المشاركون