نسخة أولية وصول مفتوح
ScopeSAE: Model-Scope Feature Discovery with Interpretable Layer Selection
Sparse autoencoders (SAEs) are a central tool in mechanistic interpretability. However, existing SAEs are primarily trained per layer. The modeling subspace is therefore fixed by layer identity, independent of which token-layer states actually drive each prediction. We argue that this constraint contributes to several …