Abstract

Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zahraei, P. S., Singh, J., Tur, G., & Hakkani-Tur, D. (2026). Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness. https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness

MLA 9

Zahraei, Pardis Sadat, et al. "Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness." https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness.

Chicago (author–date)

Zahraei, Pardis Sadat, Janvijay Singh, Gokhan Tur, and Dilek Hakkani-Tur. 2026. "Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness." https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness.

Harvard

Zahraei, P. S., Singh, J., Tur, G. and Hakkani-Tur, D. (2026) 'Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness', Available at: https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness.

Vancouver

Zahraei PS, Singh J, Tur G, Hakkani-Tur D. Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness. https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness

IEEE

P. S. Zahraei, J. Singh, G. Tur, and D. Hakkani-Tur, "Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness," https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness.