الباحثون

Eik Reichmann

المنشورات 1

نسخة أولية وصول مفتوح

Backdooring Sparse Autoencoders

Enrico Ahlers, Daniel Passon, Tobias Kiecker وآخرون · 2026

Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unch …

المؤلفون المشاركون