[
    {
        "id": "osp-23867",
        "type": "article-journal",
        "title": "Test-Time Unlearning via Sparse Autoencoder",
        "author": [
            {
                "family": "Li",
                "given": "Pingzhi"
            },
            {
                "family": "Duan",
                "given": "Jinhao"
            },
            {
                "family": "Tadiparthi",
                "given": "Vaishnav"
            },
            {
                "family": "Agarwal",
                "given": "Nakul"
            },
            {
                "family": "Lee",
                "given": "Kwonjoon"
            },
            {
                "family": "Pari",
                "given": "Ehsan Moradi"
            },
            {
                "family": "Mahjoub",
                "given": "Hossein Nourkhiz"
            },
            {
                "family": "Liu",
                "given": "Sijia"
            },
            {
                "family": "Chen",
                "given": "Tianlong"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/test-time-unlearning-via-sparse-autoencoder",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify model weights via gradient ascent and its advances. While effective on certain benchmarks, these weight-based approaches exhibit a sharp forget-utility trade-off, where stronger forgetting of target knowledge can degrade model utility, and unlearned knowledge may reappear under post-unlearning fine-tuning or prompt attacks. We propose ARIA (autoencoder-gated inference-time unlearning), a test-time unlearning method that leaves model weights intact and gates access to unwanted knowledge only when generation enters a forget-related state. ARIA uses sparse autoencoder (SAE) latents to train a lightweight linear detector, then applies an interpretable intervention on triggered states with negligible test-time overhead. Empirical evaluations on TOFU, R-TOFU, and WMDP show that ARIA improves the forget-retain trade-off over weight-based baselines across both a thinking model (DeepSeek-R1-Distilled-Qwen-1.5B) and an instruction model (Gemma-3-1B-it), e.g., reducing WMDP-cyber forget-set accuracy significantly while keeping MMLU within 1% of the pre-unlearning model. We further introduce three post-unlearning adversarial attacks targeting weight-space and decoding-space recovery, and find that ARIA remains robust under all three, with forgetting changing by less than 1% under attack. A feature-level case study leveraging the interpretability of ARIA suggests that some retain degradation may reflect response styles underlying the unlearning data rather than leakage of the targeted knowledge itself, highlighting a potential source of bias in unlearning task construction."
    }
]