الباحثون

Bennett Hillenbrand

المنشورات 2

نسخة أولية وصول مفتوح

MLCommons Jailbreak Benchmark v1.0

Carsten Maple, Cagatay Yucel, Isaac Holeman وآخرون · 2026

Modern AI systems are designed to refuse hazardous requests. A jailbreak is a prompt crafted to bypass those safeguards and elicit outputs that the system would normally refuse to provide. The MLCommons Jailbreak Benchmark v1.0 provides an end-to-end methodology for evaluating the robustness of large language models to …

نسخة أولية وصول مفتوح

Agent Reliability Profiles in Financial Services

AI agents can take actions. At times, those actions can go beyond what is intended. Agent reliability can be defined as assurance that an agent will stay within intended bounds and operate within limits. Today, there is no shared framework or language for describing, validating, and benchmarking the reliability of agen …

المؤلفون المشاركون