الباحثون

Hadi Masoudi

المنشورات 1

نسخة أولية وصول مفتوح

Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue

Omar Sheta, Rinku Deuja, Hadi Masoudi وآخرون · 2026

Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the alignment between its induced …

المؤلفون المشاركون