Abstract

Lie detection probes aim to predict from a language model's internal states whether its output is truthful or dishonest. However, role-play complicates what "truth" means for an LLM: language models can adopt a wide range of personas that take very different claims to be true, including personas whose beliefs clearly contradict reality, such as a conspiracy theorist. In this work, we investigate whether lie detection probes reliably flag falsehoods generated under such an anti-factual persona or whether they instead follow the persona's beliefs. We introduce a dataset of 8,916 human-reviewed, on-policy responses from three LLMs adopting anti-factual personas. Evaluating eight probes from prior work, we find that many fail in this setting, particularly when correct and incorrect answers are evaluated under the same persona prompt. To investigate why, we construct three novel confounder datasets in which truth is anti-correlated with a potential confounding concept. Our experiments reveal that many existing probes strongly track concepts that are spuriously correlated with truth in their training data, such as instruction compliance or response likelihood. Based on these findings, we introduce a simple linear probe that achieves the strongest overall performance on both the persona and confounder stress tests. Our results suggest that current lie detection probes are far from reliable and highlight the need for training data in which truth is decorrelated from confounding concepts.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

von Klinski, M., Lapuschkin, S., Samek, W., & Bürger, L. (2026). Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations. https://omanscience.com/en/articles/stress-testing-llm-lie-detectors-role-play-failures-and-spurious-correlations

MLA 9

von Klinski, Maximilian, et al. "Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations." https://omanscience.com/en/articles/stress-testing-llm-lie-detectors-role-play-failures-and-spurious-correlations.

Chicago (author–date)

von Klinski, Maximilian, Sebastian Lapuschkin, Wojciech Samek, and Lennart Bürger. 2026. "Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations." https://omanscience.com/en/articles/stress-testing-llm-lie-detectors-role-play-failures-and-spurious-correlations.

Harvard

von Klinski, M., Lapuschkin, S., Samek, W. and Bürger, L. (2026) 'Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations', Available at: https://omanscience.com/en/articles/stress-testing-llm-lie-detectors-role-play-failures-and-spurious-correlations.

Vancouver

von Klinski M, Lapuschkin S, Samek W, Bürger L. Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations. https://omanscience.com/en/articles/stress-testing-llm-lie-detectors-role-play-failures-and-spurious-correlations

IEEE

M. von Klinski, S. Lapuschkin, W. Samek, and L. Bürger, "Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations," https://omanscience.com/en/articles/stress-testing-llm-lie-detectors-role-play-failures-and-spurious-correlations.