الملخص

Large vision-language models (LVLMs) are increasingly deployed in safety-critical applications, yet they remain vulnerable to backdoor attacks. Defending against such attacks remains costly, as existing methods require either extensive retraining on clean data or per-query intervention at inference time. To address this limitation, we propose OrthoPurify, a more efficient method to purify backdoored model weights via one-step orthogonal projection. Specifically, through structural analysis of backdoor weight updates, we find that the backdoor is encoded by diverting a small number of weight update directions from task adaptation to backdoor shortcut encoding, a phenomenon we term direction hijacking. However, identifying these hijacked directions requires a benign reference model, which is typically inaccessible to the defender. We show that a pseudo-benign model, obtained by fine-tuning the pretrained weights on only a small set of clean samples, provides a sufficient approximation, as the dominant update directions stabilize within the first few gradient steps. OrthoPurify uses this pseudo-benign reference to isolate the hijacked directions and removes them through a single projection on the weight update. Extensive experiments show that OrthoPurify reduces the attack success rate to near zero while preserving the original performance across diverse benchmarks, without retraining the backdoored model or introducing inference-time overhead. Our code is publicly available at https://github.com/womeimingzi/OrthoPurify.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Yang, B., Zhou, H., Zhang, Z., Wang, H., Li, S., & Feng, L. (2026). Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions. https://omanscience.com/ar/articles/purifying-backdoored-large-vision-language-models-by-removing-hijacked-directions

MLA 9

Yang, Bojun, et al. "Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions." https://omanscience.com/ar/articles/purifying-backdoored-large-vision-language-models-by-removing-hijacked-directions.

شيكاغو (المؤلف–التاريخ)

Yang, Bojun, Haochen Zhou, Zhifang Zhang, Haobo Wang, Songze Li, and Lei Feng. 2026. "Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions." https://omanscience.com/ar/articles/purifying-backdoored-large-vision-language-models-by-removing-hijacked-directions.

هارفارد

Yang, B., Zhou, H., Zhang, Z., Wang, H., Li, S. and Feng, L. (2026) 'Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions', Available at: https://omanscience.com/ar/articles/purifying-backdoored-large-vision-language-models-by-removing-hijacked-directions.

فانكوفر

Yang B, Zhou H, Zhang Z, Wang H, Li S, Feng L. Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions. https://omanscience.com/ar/articles/purifying-backdoored-large-vision-language-models-by-removing-hijacked-directions

IEEE

B. Yang, H. Zhou, Z. Zhang, H. Wang, S. Li, and L. Feng, "Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions," https://omanscience.com/ar/articles/purifying-backdoored-large-vision-language-models-by-removing-hijacked-directions.