Abstract

Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, transfer slack, and leakage, and characterizes contraction relative to a base-policy risk budget evaluated on the trained policy's contexts. TRACE (Trajectory Return Attribution and Contrastive Erasure) turns this principle into a token-level objective. On the safe response, each token is weighted by the discounted return of a refusal-attributable advantage. The advantage compares a frozen reference model with its refusal-ablated copy, allowing earlier response tokens to receive credit from later refusal-related evidence. At high-gap positions on rejected responses, TRACE combines the observed token with policy-selected alternatives in the erasure target. A gradient-norm penalty replaces the retain set. Across five open-weight models and seven multi-turn attacks, TRACE gives the lowest attack success rate (ASR) in all 35 model and attack pairs, while the model utility evaluated on MMLU and HellaSwag drop by at most 1\.23 points. Source code can be found in the supplemental material.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, F., Li, K., Wang, Q., Wu, H., Zhou, J., & Wang, D. (2026). TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety. https://omanscience.com/en/articles/trace-trajectory-return-attribution-and-contrastive-erasure-for-multi-turn-safety

MLA 9

Li, Fengpeng, et al. "TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety." https://omanscience.com/en/articles/trace-trajectory-return-attribution-and-contrastive-erasure-for-multi-turn-safety.

Chicago (author–date)

Li, Fengpeng, Kemou Li, Qizhou Wang, Haiwei Wu, Jiantao Zhou, and Di Wang. 2026. "TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety." https://omanscience.com/en/articles/trace-trajectory-return-attribution-and-contrastive-erasure-for-multi-turn-safety.

Harvard

Li, F., Li, K., Wang, Q., Wu, H., Zhou, J. and Wang, D. (2026) 'TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety', Available at: https://omanscience.com/en/articles/trace-trajectory-return-attribution-and-contrastive-erasure-for-multi-turn-safety.

Vancouver

Li F, Li K, Wang Q, Wu H, Zhou J, Wang D. TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety. https://omanscience.com/en/articles/trace-trajectory-return-attribution-and-contrastive-erasure-for-multi-turn-safety

IEEE

F. Li, K. Li, Q. Wang, H. Wu, J. Zhou, and D. Wang, "TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety," https://omanscience.com/en/articles/trace-trajectory-return-attribution-and-contrastive-erasure-for-multi-turn-safety.