الملخص

Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference. However, their ability to generate harmful or undesired content poses significant safety risks. Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives. Crucially, these objectives implicitly rely on multi-step denoising dynamics, an assumption that breaks down for few-step distilled (FSD) models, resulting in ineffective forgetting. Furthermore, performing unlearning on the non-distilled base model and subsequently re-distilling it to obtain an unlearned FSD model incurs substantial computational and time overhead, making it impractical in many settings. Hence, we address this limitation with a preference-driven unlearning framework that revisits Direct Preference Optimization (DPO) for diffusion models. We show that standard DPO and its unlearning derivatives, formulated around noise-prediction error, transfer poorly to FSD models due to their altered generation dynamics. To overcome this, we introduce a modified preference optimization formulation explicitly aligned with the few-step generation properties, enabling direct concept removal in FSD models while preserving few-step efficiency and maintaining strong retention of desirable (non-targeted) capabilities. We evaluate our framework primarily on identity and NSFW (nudity) removal tasks and also extend our method to object-level unlearning. Extensive experiments demonstrate consistent and effective forgetting, and strong retention performance, establishing our method as a practical and principled solution for unlearning in FSD models.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Patel, G., Fang, J., Steeg, G. V., Qiu, Q., & Sripada, S. (2026). Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models. https://omanscience.com/ar/articles/enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-diffusion-models

MLA 9

Patel, Gaurav, et al. "Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models." https://omanscience.com/ar/articles/enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-diffusion-models.

شيكاغو (المؤلف–التاريخ)

Patel, Gaurav, Jun Fang, Greg Ver Steeg, Qiang Qiu, and Sravan Sripada. 2026. "Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models." https://omanscience.com/ar/articles/enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-diffusion-models.

هارفارد

Patel, G., Fang, J., Steeg, G. V., Qiu, Q. and Sripada, S. (2026) 'Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models', Available at: https://omanscience.com/ar/articles/enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-diffusion-models.

فانكوفر

Patel G, Fang J, Steeg GV, Qiu Q, Sripada S. Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models. https://omanscience.com/ar/articles/enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-diffusion-models

IEEE

G. Patel, J. Fang, G. V. Steeg, Q. Qiu, and S. Sripada, "Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models," https://omanscience.com/ar/articles/enabling-preference-driven-unlearning-in-few-step-distilled-text-to-image-diffusion-models.