الملخص
LoRAs are widely studied for adapting base text-to-image diffusion models. However, a backdoored LoRA can hide a backdoor: it behaves normally in most cases, but produces attacker-specified content (the backdoor target) when a hidden backdoor trigger appears in the input prompt. Detecting such backdoors before using an untrusted LoRA is important for the safety of LoRA adaptation. We present TokenScanner, a model-level vocabulary scanner for backdoor detection within a LoRA fine-tuned text-to-image diffusion model, aiming to discover the malicious trigger for trustworthy LoRA adaptation. The key observation is that backdoor trigger tokens that appear in a larger proportion of training prompts tend to induce more prominent token-specific LoRA responses than those induced by unrelated tokens. TokenScanner therefore scans the tokenizer vocabulary and measures token-wise LoRA responses in the U-Net and the text encoder. It uses these responses to detect backdoored LoRAs and rank candidate trigger tokens for subsequent testing of backdoor activation. Experiments on seven backdoor settings, comprising 840 backdoored LoRAs and 840 real-world benign test LoRAs, show that TokenScanner achieves 95.96% AUC and 96.55% TPR at an FPR of 10.95%. It also achieves 89.40% Hit@1 and 99.40% Hit@5 for trigger discovery across all seven settings.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Liu, B., & Zhang, J. (2026). TokenScanner: Detecting Backdoors and Discovering Triggers in Text-to-Image LoRAs via Full Vocabulary Scanning. https://omanscience.com/ar/articles/tokenscanner-detecting-backdoors-and-discovering-triggers-in-text-to-image-loras-via-full-vocabulary-scanning
MLA 9
Liu, Boliang, and Jing Zhang. "TokenScanner: Detecting Backdoors and Discovering Triggers in Text-to-Image LoRAs via Full Vocabulary Scanning." https://omanscience.com/ar/articles/tokenscanner-detecting-backdoors-and-discovering-triggers-in-text-to-image-loras-via-full-vocabulary-scanning.
شيكاغو (المؤلف–التاريخ)
Liu, Boliang, and Jing Zhang. 2026. "TokenScanner: Detecting Backdoors and Discovering Triggers in Text-to-Image LoRAs via Full Vocabulary Scanning." https://omanscience.com/ar/articles/tokenscanner-detecting-backdoors-and-discovering-triggers-in-text-to-image-loras-via-full-vocabulary-scanning.
هارفارد
Liu, B. and Zhang, J. (2026) 'TokenScanner: Detecting Backdoors and Discovering Triggers in Text-to-Image LoRAs via Full Vocabulary Scanning', Available at: https://omanscience.com/ar/articles/tokenscanner-detecting-backdoors-and-discovering-triggers-in-text-to-image-loras-via-full-vocabulary-scanning.
فانكوفر
Liu B, Zhang J. TokenScanner: Detecting Backdoors and Discovering Triggers in Text-to-Image LoRAs via Full Vocabulary Scanning. https://omanscience.com/ar/articles/tokenscanner-detecting-backdoors-and-discovering-triggers-in-text-to-image-loras-via-full-vocabulary-scanning
IEEE
B. Liu, and J. Zhang, "TokenScanner: Detecting Backdoors and Discovering Triggers in Text-to-Image LoRAs via Full Vocabulary Scanning," https://omanscience.com/ar/articles/tokenscanner-detecting-backdoors-and-discovering-triggers-in-text-to-image-loras-via-full-vocabulary-scanning.