نسخة أولية وصول مفتوح
Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models
Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a primary component can inhibit the activation of a backup, leading to issues with ranking components. Actual causality studies …