الباحثون

Xujie Si

المنشورات 2

نسخة أولية وصول مفتوح

Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

Chuqin Geng, Li Zhang, Haolin Ye وآخرون · 2026

Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circ …

نسخة أولية وصول مفتوح

Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?

Li Zhang, Chuqin Geng, Mark Zhang وآخرون · 2026

Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlyin …

المؤلفون المشاركون