Abstract

Open-weight LLMs are released not only as fixed products but also as substrates for downstream fine-tuning. This openness, however, creates legal and ethical risks because users may misuse fine-tuning to instill illicit knowledge or enable hostile operations. Model providers therefore need apre-release defense against such acquisition, motivating the problem of preemptive unlearning. Unlike retrospective unlearning, which removes capabilities already present in a fixed model, preemptive unlearning seeks to prevent their acquisition under unseen attack data and future fine-tuning procedures. Despite its practical importance, this setting remains largely unexplored, presents distinct challenges, and is therefore the central focus of our work. We first verify that existing retrospective methods provide insufficient pre-release protection. Even when forbidden capabilities are suppressed in current outputs, forbidden-domain data can still induce gradients through internal pathways, enabling later acquisition. Motivated by this finding, we propose a gradient-sealing principle that blocks these pathways by pushing relevant pre-activations into the negative region, where ReLU-family activations exhibit zero or near-zero derivatives. Experiments across multiple LLM families demonstrate our stronger resistance to downstream acquisition than retrospective baselines, validating gradient sealing as an effective mechanism for pre-release protection.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, K., Wang, Q., Wang, Y., Li, F., Shi, Z., Rostamzadeh, N., Farnadi, G., Sugiyama, M., & Zhou, J. (2026). Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing. https://omanscience.com/en/articles/preemptive-llm-unlearning-against-forbidden-capability-acquisition-via-gradient-sealing

MLA 9

Li, Kemou, et al. "Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing." https://omanscience.com/en/articles/preemptive-llm-unlearning-against-forbidden-capability-acquisition-via-gradient-sealing.

Chicago (author–date)

Li, Kemou, Qizhou Wang, Yue Wang, Fengpeng Li, Zhuan Shi, Negar Rostamzadeh, Golnoosh Farnadi, Masashi Sugiyama, and Jiantao Zhou. 2026. "Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing." https://omanscience.com/en/articles/preemptive-llm-unlearning-against-forbidden-capability-acquisition-via-gradient-sealing.

Harvard

Li, K., Wang, Q., Wang, Y., Li, F., Shi, Z., Rostamzadeh, N., Farnadi, G., Sugiyama, M. and Zhou, J. (2026) 'Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing', Available at: https://omanscience.com/en/articles/preemptive-llm-unlearning-against-forbidden-capability-acquisition-via-gradient-sealing.

Vancouver

Li K, Wang Q, Wang Y, Li F, Shi Z, Rostamzadeh N, et al. Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing. https://omanscience.com/en/articles/preemptive-llm-unlearning-against-forbidden-capability-acquisition-via-gradient-sealing

IEEE

K. Li, Q. Wang, Y. Wang, F. Li, Z. Shi, N. Rostamzadeh, G. Farnadi, M. Sugiyama, and J. Zhou, "Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing," https://omanscience.com/en/articles/preemptive-llm-unlearning-against-forbidden-capability-acquisition-via-gradient-sealing.