Abstract

Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints. To address this, we propose SafeEvo, an interpretability framework from the circuit (sparse subgraphs of an LLM) perspective. SafeEvo first applies an optimization-based extraction algorithm to identify weak refusal circuits in pretrained base LLMs that can independently express refusal behavior. Causally ablating these circuits completely eliminates the base model's refusal of harmful inputs. SafeEvo then traces the evolution of refusal circuits across successive alignment checkpoints and finds that their structures change progressively, suggesting that the alignment tax may result from refusal-circuit updates affecting utility-related parameters. To validate this, SafeEvo introduces Safety Circuit Alignment (SCA), which confines safety updates to the refusal circuits. Experiments across three LLMs and two alignment algorithms show that, on average, SCA outperforms vanilla alignment in three aspects: \textbf{(1) stronger alignment}, lowering harmfulness score by 63.21\%; \textbf{(2) less over-refusal}, yielding a 58.44\% decrease in refusal rates for benign queries; and \textbf{(3) better utility}, retaining 99.58\% of the original model capabilities.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Yu, M., Huang, H., Yuan, L., Li, Y., Wang, K., & Jiang, Z. (2026). SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models. https://omanscience.com/en/articles/safeevo-deciphering-the-safety-alignment-mechanism-and-evolution-in-language-models

MLA 9

Yu, Miao, et al. "SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models." https://omanscience.com/en/articles/safeevo-deciphering-the-safety-alignment-mechanism-and-evolution-in-language-models.

Chicago (author–date)

Yu, Miao, Hao Huang, Lu Yuan, Yunpeng Li, Kun Wang, and Zuming Jiang. 2026. "SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models." https://omanscience.com/en/articles/safeevo-deciphering-the-safety-alignment-mechanism-and-evolution-in-language-models.

Harvard

Yu, M., Huang, H., Yuan, L., Li, Y., Wang, K. and Jiang, Z. (2026) 'SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models', Available at: https://omanscience.com/en/articles/safeevo-deciphering-the-safety-alignment-mechanism-and-evolution-in-language-models.

Vancouver

Yu M, Huang H, Yuan L, Li Y, Wang K, Jiang Z. SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models. https://omanscience.com/en/articles/safeevo-deciphering-the-safety-alignment-mechanism-and-evolution-in-language-models

IEEE

M. Yu, H. Huang, L. Yuan, Y. Li, K. Wang, and Z. Jiang, "SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models," https://omanscience.com/en/articles/safeevo-deciphering-the-safety-alignment-mechanism-and-evolution-in-language-models.