Abstract

Can reasoning models trick chain of thought (CoT) monitors and perform hidden computation without revealing it in their thinking traces? We show that the answer depends on the underlying task difficulty and the model size. Simple computations can be performed covertly; however, beyond a threshold depending on model size, successfully solving the task necessarily leaks a near-linear amount of information about the covert task input into the CoT. Therefore, sufficiently complex hidden computation always leaves an information-theoretic footprint. However, concerningly, this leakage need not be readable: Under plausible cryptographic assumptions, even a one-layer Transformer can encrypt its reasoning online so that no polynomial-time monitor can extract information about the hidden computation. Overall, our theoretical and empirical results provide a holistic view of both the opportunities and the limitations of CoT monitoring.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Mohammadkhani, M., Krishna, M., Sarrof, Y., & Hahn, M. (2026). Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring. https://omanscience.com/en/articles/hidden-reasoning-must-leak-but-need-not-be-readable-fundamental-opportunities-and-limits-for-chain-of-thought-monitoring

MLA 9

Mohammadkhani, Mohammadali, et al. "Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring." https://omanscience.com/en/articles/hidden-reasoning-must-leak-but-need-not-be-readable-fundamental-opportunities-and-limits-for-chain-of-thought-monitoring.

Chicago (author–date)

Mohammadkhani, Mohammadali, Madhava Krishna, Yash Sarrof, and Michael Hahn. 2026. "Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring." https://omanscience.com/en/articles/hidden-reasoning-must-leak-but-need-not-be-readable-fundamental-opportunities-and-limits-for-chain-of-thought-monitoring.

Harvard

Mohammadkhani, M., Krishna, M., Sarrof, Y. and Hahn, M. (2026) 'Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring', Available at: https://omanscience.com/en/articles/hidden-reasoning-must-leak-but-need-not-be-readable-fundamental-opportunities-and-limits-for-chain-of-thought-monitoring.

Vancouver

Mohammadkhani M, Krishna M, Sarrof Y, Hahn M. Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring. https://omanscience.com/en/articles/hidden-reasoning-must-leak-but-need-not-be-readable-fundamental-opportunities-and-limits-for-chain-of-thought-monitoring

IEEE

M. Mohammadkhani, M. Krishna, Y. Sarrof, and M. Hahn, "Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring," https://omanscience.com/en/articles/hidden-reasoning-must-leak-but-need-not-be-readable-fundamental-opportunities-and-limits-for-chain-of-thought-monitoring.