الملخص

Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model's alignment. Two recent alignment-stage defenses address this problem at different levels of the model. Vaccine improves the robustness of hidden embeddings to the representation shifts induced by harmful fine-tuning, whereas Booster simulates harmful weight updates and attenuates their effect during alignment. We investigate whether these mechanisms are complementary and propose VaccineBooster, a single alignment procedure that combines embedding perturbation and weight-level gradient attenuation within each training step. On Llama-2-7B aligned with BeaverTails and then attacked through poisoned fine-tuning, VaccineBooster achieves the lowest OpenAI moderation score among the compared defenses, 0.315, while a Booster-Only variant retains the highest post-attack refusal rate, 50%. Together with ablations over the embedding-perturbation and gradient-attenuation strengths, these results indicate a trade-off: embedding perturbation primarily reduces flagged harmful content, whereas gradient attenuation primarily preserves explicit refusal behavior. Because our evaluation uses ten prompts and a single unseeded run per configuration, we report this trade-off as an observed pattern rather than a statistically resolved effect. These results provide practical guidance for prioritizing content safety or refusal retention when aligned models are exposed to untrusted fine-tuning.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Akram, M. Z., Marican, M. K., Yenugu, A. R., Kaimkhani, A. Z., & Fang, M. (2026). Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning. https://omanscience.com/ar/articles/safer-content-or-firmer-refusals-a-hybrid-perturbation-defense-for-alignment-under-harmful-fine-tuning

MLA 9

Akram, Muhammad Zeeshan, et al. "Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning." https://omanscience.com/ar/articles/safer-content-or-firmer-refusals-a-hybrid-perturbation-defense-for-alignment-under-harmful-fine-tuning.

شيكاغو (المؤلف–التاريخ)

Akram, Muhammad Zeeshan, Mufid Kamel Marican, Anvesh Reddy Yenugu, Ali Zain Kaimkhani, and Minghong Fang. 2026. "Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning." https://omanscience.com/ar/articles/safer-content-or-firmer-refusals-a-hybrid-perturbation-defense-for-alignment-under-harmful-fine-tuning.

هارفارد

Akram, M. Z., Marican, M. K., Yenugu, A. R., Kaimkhani, A. Z. and Fang, M. (2026) 'Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning', Available at: https://omanscience.com/ar/articles/safer-content-or-firmer-refusals-a-hybrid-perturbation-defense-for-alignment-under-harmful-fine-tuning.

فانكوفر

Akram MZ, Marican MK, Yenugu AR, Kaimkhani AZ, Fang M. Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning. https://omanscience.com/ar/articles/safer-content-or-firmer-refusals-a-hybrid-perturbation-defense-for-alignment-under-harmful-fine-tuning

IEEE

M. Z. Akram, M. K. Marican, A. R. Yenugu, A. Z. Kaimkhani, and M. Fang, "Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning," https://omanscience.com/ar/articles/safer-content-or-firmer-refusals-a-hybrid-perturbation-defense-for-alignment-under-harmful-fine-tuning.