الباحثون

Wenxuan Ding

المنشورات 1

نسخة أولية وصول مفتوح

Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models

Ziyuan Yang, Wenxuan Ding, Shangbin Feng وآخرون · 2026

Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing- …

المؤلفون المشاركون