Abstract

Alignment does not eliminate behavioral errors in language models. Models may still refuse benign requests, call unnecessary tools, or yield to false user claims. Current methods mitigate such errors as a computation problem, and rarely explore if the desired behavior is already encoded in the model's representation. Motivated by the observation that behavior-relevant information remains linearly decodable from the final hidden state even when the resulting logits produce the undesired behavior, we introduce HeadEdit, a gradient-free method that calibrates model behavior through the unembedding matrix. HeadEdit extracts a low-rank behavioral subspace from paired completions and uses each prompt's coordinates within it to generate a vocabulary-wide correction, thereby implementing implicitly adaptive steering without manually specified target tokens or parameter updates. HeadEdit improves all nine experimental settings across three tasks and three model families, with negligible inference overhead and no systematic loss of general capabilities. It also reveals a connection to gradient-based alignment. HeadEdit's low-dimensional representation partly predicts how preference tuning changes output logits on unseen prompts. The subspace learned from the model can also be reused after tuning, improving performance without re-extracting or retuning. These results show that HeadEdit provides a practical, lightweight, and interpretable way to calibrate model behavior through the unembedding matrix.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

He, Z., Zhao, H., Hu, J., Wu, Y., Yuan, C., Li, Y., Bai, Y., & Du, M. (2026). HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix. https://omanscience.com/en/articles/headedit-calibrating-language-model-behavior-through-the-frozen-unembedding-matrix

MLA 9

He, Zirui, et al. "HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix." https://omanscience.com/en/articles/headedit-calibrating-language-model-behavior-through-the-frozen-unembedding-matrix.

Chicago (author–date)

He, Zirui, Haiyan Zhao, Jingyu Hu, Yinghao Wu, Chenxi Yuan, Yingcong Li, Yandong Bai, and Mengnan Du. 2026. "HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix." https://omanscience.com/en/articles/headedit-calibrating-language-model-behavior-through-the-frozen-unembedding-matrix.

Harvard

He, Z., Zhao, H., Hu, J., Wu, Y., Yuan, C., Li, Y., Bai, Y. and Du, M. (2026) 'HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix', Available at: https://omanscience.com/en/articles/headedit-calibrating-language-model-behavior-through-the-frozen-unembedding-matrix.

Vancouver

He Z, Zhao H, Hu J, Wu Y, Yuan C, Li Y, et al. HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix. https://omanscience.com/en/articles/headedit-calibrating-language-model-behavior-through-the-frozen-unembedding-matrix

IEEE

Z. He, H. Zhao, J. Hu, Y. Wu, C. Yuan, Y. Li, Y. Bai, and M. Du, "HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix," https://omanscience.com/en/articles/headedit-calibrating-language-model-behavior-through-the-frozen-unembedding-matrix.