نسخة أولية وصول مفتوح
Component and Dimension Sparsity in Transformer Refusal Mechanisms
Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly understood. We decompose refusal steering into component-level interventions across four open-weight models, identifying the sparse subsets of attention a …