الملخص

Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify model weights via gradient ascent and its advances. While effective on certain benchmarks, these weight-based approaches exhibit a sharp forget-utility trade-off, where stronger forgetting of target knowledge can degrade model utility, and unlearned knowledge may reappear under post-unlearning fine-tuning or prompt attacks. We propose ARIA (autoencoder-gated inference-time unlearning), a test-time unlearning method that leaves model weights intact and gates access to unwanted knowledge only when generation enters a forget-related state. ARIA uses sparse autoencoder (SAE) latents to train a lightweight linear detector, then applies an interpretable intervention on triggered states with negligible test-time overhead. Empirical evaluations on TOFU, R-TOFU, and WMDP show that ARIA improves the forget-retain trade-off over weight-based baselines across both a thinking model (DeepSeek-R1-Distilled-Qwen-1.5B) and an instruction model (Gemma-3-1B-it), e.g., reducing WMDP-cyber forget-set accuracy significantly while keeping MMLU within 1% of the pre-unlearning model. We further introduce three post-unlearning adversarial attacks targeting weight-space and decoding-space recovery, and find that ARIA remains robust under all three, with forgetting changing by less than 1% under attack. A feature-level case study leveraging the interpretability of ARIA suggests that some retain degradation may reflect response styles underlying the unlearning data rather than leakage of the targeted knowledge itself, highlighting a potential source of bias in unlearning task construction.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Li, P., Duan, J., Tadiparthi, V., Agarwal, N., Lee, K., Pari, E. M., Mahjoub, H. N., Liu, S., & Chen, T. (2026). Test-Time Unlearning via Sparse Autoencoder. https://omanscience.com/ar/articles/test-time-unlearning-via-sparse-autoencoder

MLA 9

Li, Pingzhi, et al. "Test-Time Unlearning via Sparse Autoencoder." https://omanscience.com/ar/articles/test-time-unlearning-via-sparse-autoencoder.

شيكاغو (المؤلف–التاريخ)

Li, Pingzhi, Jinhao Duan, Vaishnav Tadiparthi, Nakul Agarwal, Kwonjoon Lee, Ehsan Moradi Pari, Hossein Nourkhiz Mahjoub, Sijia Liu, and Tianlong Chen. 2026. "Test-Time Unlearning via Sparse Autoencoder." https://omanscience.com/ar/articles/test-time-unlearning-via-sparse-autoencoder.

هارفارد

Li, P., Duan, J., Tadiparthi, V., Agarwal, N., Lee, K., Pari, E. M., Mahjoub, H. N., Liu, S. and Chen, T. (2026) 'Test-Time Unlearning via Sparse Autoencoder', Available at: https://omanscience.com/ar/articles/test-time-unlearning-via-sparse-autoencoder.

فانكوفر

Li P, Duan J, Tadiparthi V, Agarwal N, Lee K, Pari EM, et al. Test-Time Unlearning via Sparse Autoencoder. https://omanscience.com/ar/articles/test-time-unlearning-via-sparse-autoencoder

IEEE

P. Li, J. Duan, V. Tadiparthi, N. Agarwal, K. Lee, E. M. Pari, H. N. Mahjoub, S. Liu, and T. Chen, "Test-Time Unlearning via Sparse Autoencoder," https://omanscience.com/ar/articles/test-time-unlearning-via-sparse-autoencoder.