Abstract

Recent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people. We show that white-box deception detection via probes can be scaled up to frontier monitoring settings by collecting the largest deception dataset to date for training probes and introducing a novel probe architecture which can aggregate information across many layers and tokens. Our probes achieve 98.8% AUC in SHADE-Arena, surpassing an Opus 5.5 text-monitoring baseline, and show improved efficacy as the underlying model is scaled up. To push our probes to their limit, we test them on several cases where deception cannot be determined from the context alone. In these cases, which we refer to as introspective deception, the ground truth can only be determined through careful elicitation or thorough knowledge of a model's training data. In one such evaluation, we show that probes can distinguish transcripts containing a model's true hidden goal from other goals with an AUC of up to 99.7%. Our probes also readily detect deception on prominent open-weight models which lie about politically sensitive topics, and about their beliefs when put under pressure. We release our training dataset, dubbed FIBS, to help drive frontier deployment of effective probes, and encourage the community to expand upon it with further examples of deception and sabotage.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Hollinsworth, O. J., Spies, A. F., Diriba, T., Gleave, A., & Cundy, C. (2026). Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception. https://omanscience.com/en/articles/caught-in-the-act-probes-effectively-detect-sabotage-and-catch-unverbalized-deception

MLA 9

Hollinsworth, Oskar J., et al. "Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception." https://omanscience.com/en/articles/caught-in-the-act-probes-effectively-detect-sabotage-and-catch-unverbalized-deception.

Chicago (author–date)

Hollinsworth, Oskar J., Alex F. Spies, Tigist Diriba, Adam Gleave, and Chris Cundy. 2026. "Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception." https://omanscience.com/en/articles/caught-in-the-act-probes-effectively-detect-sabotage-and-catch-unverbalized-deception.

Harvard

Hollinsworth, O. J., Spies, A. F., Diriba, T., Gleave, A. and Cundy, C. (2026) 'Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception', Available at: https://omanscience.com/en/articles/caught-in-the-act-probes-effectively-detect-sabotage-and-catch-unverbalized-deception.

Vancouver

Hollinsworth OJ, Spies AF, Diriba T, Gleave A, Cundy C. Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception. https://omanscience.com/en/articles/caught-in-the-act-probes-effectively-detect-sabotage-and-catch-unverbalized-deception

IEEE

O. J. Hollinsworth, A. F. Spies, T. Diriba, A. Gleave, and C. Cundy, "Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception," https://omanscience.com/en/articles/caught-in-the-act-probes-effectively-detect-sabotage-and-catch-unverbalized-deception.