الباحثون

Francesco Croce

المنشورات 1

نسخة أولية وصول مفتوح

Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment

Shunchang Liu, Lukas Fluri, Xin Chen وآخرون · 2026

Modern AI models are aligned through post-training to adapt them to downstream tasks. Recent work shows that fine-tuning language models on narrow tasks can induce emergent misalignment (EM), causing broadly harmful behaviors beyond the training task. However, EM has been studied almost entirely in text-only tasks, lea …

المؤلفون المشاركون