Preprint Open access
Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it
Unlearning in large language models (LLMs) is typically evaluated at the output level, where a model appears to suppress sensitive or undesirable content. In this work, we show that such evaluations can create an illusion of forgetting: even when output-level leakage is eliminated, sensitive information can remain enco …