الباحثون

Jonathan Petit

المنشورات 2

نسخة أولية وصول مفتوح

MLCommons Jailbreak Benchmark v1.0

Carsten Maple, Cagatay Yucel, Isaac Holeman وآخرون · 2026

Modern AI systems are designed to refuse hazardous requests. A jailbreak is a prompt crafted to bypass those safeguards and elicit outputs that the system would normally refuse to provide. The MLCommons Jailbreak Benchmark v1.0 provides an end-to-end methodology for evaluating the robustness of large language models to …

نسخة أولية وصول مفتوح

Beyond Predictable Paths: Redefining AI Security Incident Reporting for Agents

AI agents are being deployed rapidly, accompanied by a growing number of AI-specific attacks and corresponding incidents. As incident reporting becomes increasingly important for legal compliance, governance, accountability, and security; current frameworks must be adapted to the unique characteristics of AI agents. In …

المؤلفون المشاركون