Abstract

Reinforcement learning from human feedback (RLHF) has played a central role in making large language models responsive to human instructions. However, human evaluators often favor flattering or persuasive responses over truthful ones, creating incentives for models to appeal to evaluators at the expense of accuracy. AI safety via debate has been proposed as a way to improve the supervision of language models: in this paradigm, two agents argue opposing positions and challenge each other's claims, potentially exposing falsehoods to the adjudicator. A central premise of AI safety via debate is that truthful arguments are easier to defend than false ones under adversarial scrutiny. In this work, we investigate whether this advantage persists when debaters use rhetorical strategies that exploit biases in human judgment. Inspired by competitive debate, we construct 68 LLM-generated dialogues about detective mysteries with known culprits, spanning four interventions: anchoring, fallacy oversight, pro-jargon, and verbosity. We apply each intervention to either the side advocating for the true culprit or the side advocating for an innocent suspect, allowing us to distinguish influence on adjudication from correctness. In a study with 369 participants, we find that, pooled across bias types, these interventions significantly shift judgments toward the manipulated side. These findings expose a vulnerability in debate-based supervision: human adjudication is sensitive to manipulative rhetorical strategies.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Liu, G., Rashkovan, S., George, S. L., Sheidlower, I., & Booth, S. (2026). AI Safety via Debate is Compromised by Cognitive Biases. https://omanscience.com/en/articles/ai-safety-via-debate-is-compromised-by-cognitive-biases

MLA 9

Liu, Gefei, et al. "AI Safety via Debate is Compromised by Cognitive Biases." https://omanscience.com/en/articles/ai-safety-via-debate-is-compromised-by-cognitive-biases.

Chicago (author–date)

Liu, Gefei, Sonya Rashkovan, Sophia Lloyd George, Isaac Sheidlower, and Serena Booth. 2026. "AI Safety via Debate is Compromised by Cognitive Biases." https://omanscience.com/en/articles/ai-safety-via-debate-is-compromised-by-cognitive-biases.

Harvard

Liu, G., Rashkovan, S., George, S. L., Sheidlower, I. and Booth, S. (2026) 'AI Safety via Debate is Compromised by Cognitive Biases', Available at: https://omanscience.com/en/articles/ai-safety-via-debate-is-compromised-by-cognitive-biases.

Vancouver

Liu G, Rashkovan S, George SL, Sheidlower I, Booth S. AI Safety via Debate is Compromised by Cognitive Biases. https://omanscience.com/en/articles/ai-safety-via-debate-is-compromised-by-cognitive-biases

IEEE

G. Liu, S. Rashkovan, S. L. George, I. Sheidlower, and S. Booth, "AI Safety via Debate is Compromised by Cognitive Biases," https://omanscience.com/en/articles/ai-safety-via-debate-is-compromised-by-cognitive-biases.