Preprint Open access
SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models
Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment …