Abstract

Modern large language models (LLMs) are trained on massive, largely unfiltered datasets, including content scraped from nearly every accessible website and user inputs. As a result, LLMs often memorize and reproduce personally sensitive information (PSI) such as birth dates, phone numbers, and home addresses. This leads to significant privacy risks, particularly for high-profile individuals such as executives, politicians, and judges. Existing mitigations largely rely on machine unlearning. However, these methods often remove more information than needed, degrade model utility and safety, and are highly vulnerable to attacks. This paper presents Whiteout, a practical tool that, upon requests by individuals, prevents LLMs from regurgitating their genuine PSIs, by overwriting them using precise and carefully designed obfuscation samples. We evaluate Whiteout on modern LLMs of varying sizes and makers, including a widely-used OpenAI model. Results show that Whiteout effectively prevents disclosure of the targeted PSIs, has negligible impact on model utility and safety, and outperforms existing alternatives. We also test Whiteout against a wide range of countermeasures, from black-box attacks like jailbreaking to white-box adaptive attacks like relearning and quantization. Finally, we conclude with a discussion on the security and ethical implications of Whiteout.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Ha, A. Y. J., Bhaskar, R., Zheng, H., & Zhao, B. Y. (2026). Mitigating Private Data Leakage in LLMs with Whiteout. https://omanscience.com/en/articles/mitigating-private-data-leakage-in-llms-with-whiteout

MLA 9

Ha, Anna Yoo Jeong, et al. "Mitigating Private Data Leakage in LLMs with Whiteout." https://omanscience.com/en/articles/mitigating-private-data-leakage-in-llms-with-whiteout.

Chicago (author–date)

Ha, Anna Yoo Jeong, Ronik Bhaskar, Haitao Zheng, and Ben Y. Zhao. 2026. "Mitigating Private Data Leakage in LLMs with Whiteout." https://omanscience.com/en/articles/mitigating-private-data-leakage-in-llms-with-whiteout.

Harvard

Ha, A. Y. J., Bhaskar, R., Zheng, H. and Zhao, B. Y. (2026) 'Mitigating Private Data Leakage in LLMs with Whiteout', Available at: https://omanscience.com/en/articles/mitigating-private-data-leakage-in-llms-with-whiteout.

Vancouver

Ha AYJ, Bhaskar R, Zheng H, Zhao BY. Mitigating Private Data Leakage in LLMs with Whiteout. https://omanscience.com/en/articles/mitigating-private-data-leakage-in-llms-with-whiteout

IEEE

A. Y. J. Ha, R. Bhaskar, H. Zheng, and B. Y. Zhao, "Mitigating Private Data Leakage in LLMs with Whiteout," https://omanscience.com/en/articles/mitigating-private-data-leakage-in-llms-with-whiteout.