Abstract

Unlearning in large language models (LLMs) is typically evaluated at the output level, where a model appears to suppress sensitive or undesirable content. In this work, we show that such evaluations can create an illusion of forgetting: even when output-level leakage is eliminated, sensitive information can remain encoded in the model's hidden representations. We first provide a theoretical analysis establishing a fundamental separation between output suppression and representational erasure. Specifically, we show that the decoder can be made arbitrarily insensitive to sensitive directions, driving output-level leakage to zero, while the hidden representations retain the underlying information. To empirically validate this phenomenon, we train generative probe decoders on hidden states across transformer layers, enabling layer-wise measurement of information leakage. Across three widely used benchmarks, TOFU, MUSE, and WMDP, and state-of-the-art unlearning methods, we find that substantial sensitive information remains recoverable from hidden representations, even when standard output-level metrics indicate successful unlearning. To address this gap, we propose Probe-Adversarial Representation Suppression (PARS), an unlearning objective that adversarially minimizes the extractable information from hidden representations. PARS directly targets representational leakage and provides significantly stronger guarantees of erasure under adversarial probing and relearning attacks, outperforming all evaluated baselines. Our results highlight a fundamental limitation of existing unlearning paradigms and suggest that true forgetting in LLMs requires controlling not only model outputs, but also the information encoded in hidden representations. Codes are available at https://github.com/OptimAI-Lab/HiddenStateUnlearning.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Reisizadeh, H., Ruan, J., Liu, S., & Hong, M. (2026). Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it. https://omanscience.com/en/articles/do-llms-really-forget-hidden-state-leakage-in-model-unlearning-and-how-to-fix-it

MLA 9

Reisizadeh, Hadi, et al. "Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it." https://omanscience.com/en/articles/do-llms-really-forget-hidden-state-leakage-in-model-unlearning-and-how-to-fix-it.

Chicago (author–date)

Reisizadeh, Hadi, Jiajun Ruan, Sijia Liu, and Mingyi Hong. 2026. "Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it." https://omanscience.com/en/articles/do-llms-really-forget-hidden-state-leakage-in-model-unlearning-and-how-to-fix-it.

Harvard

Reisizadeh, H., Ruan, J., Liu, S. and Hong, M. (2026) 'Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it', Available at: https://omanscience.com/en/articles/do-llms-really-forget-hidden-state-leakage-in-model-unlearning-and-how-to-fix-it.

Vancouver

Reisizadeh H, Ruan J, Liu S, Hong M. Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it. https://omanscience.com/en/articles/do-llms-really-forget-hidden-state-leakage-in-model-unlearning-and-how-to-fix-it

IEEE

H. Reisizadeh, J. Ruan, S. Liu, and M. Hong, "Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it," https://omanscience.com/en/articles/do-llms-really-forget-hidden-state-leakage-in-model-unlearning-and-how-to-fix-it.