الباحثون

Haiwei Wu

المنشورات 2

نسخة أولية وصول مفتوح

TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

Fengpeng Li, Kemou Li, Qizhou Wang وآخرون · 2026

Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which …

المؤلفون المشاركون