الباحثون

Zirui He

المنشورات 1

نسخة أولية وصول مفتوح

HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix

Zirui He, Haiyan Zhao, Jingyu Hu وآخرون · 2026

Alignment does not eliminate behavioral errors in language models. Models may still refuse benign requests, call unnecessary tools, or yield to false user claims. Current methods mitigate such errors as a computation problem, and rarely explore if the desired behavior is already encoded in the model's representation. M …

المؤلفون المشاركون