Abstract
Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zahraei, P. S., Singh, J., Tur, G., & Hakkani-Tur, D. (2026). Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness. https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness
MLA 9
Zahraei, Pardis Sadat, et al. "Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness." https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness.
Chicago (author–date)
Zahraei, Pardis Sadat, Janvijay Singh, Gokhan Tur, and Dilek Hakkani-Tur. 2026. "Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness." https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness.
Harvard
Zahraei, P. S., Singh, J., Tur, G. and Hakkani-Tur, D. (2026) 'Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness', Available at: https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness.
Vancouver
Zahraei PS, Singh J, Tur G, Hakkani-Tur D. Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness. https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness
IEEE
P. S. Zahraei, J. Singh, G. Tur, and D. Hakkani-Tur, "Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness," https://omanscience.com/en/articles/emergent-unfaithfulness-how-alignment-training-causes-language-models-to-silently-override-task-faithfulness.