Preprint Open access
TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety
Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which …