الباحثون

Xiaochun Cao

المنشورات 5

نسخة أولية وصول مفتوح

Representation--Behavior Alignment for Explainable Weakly-Supervised Video Anomaly Detection

Chao Huang, Pengfei Wei, Kaige Li وآخرون · 2026

Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as representation--behavior misalignment. We decompose this gap in …

نسخة أولية وصول مفتوح

RoMod: Temporal Routing Modulation via Mixture-of-Experts for Video Anomaly Detection

Chao Huang, Pengfei Wei, Benfeng Wang وآخرون · 2026

Intermediate-layer features from multimodal large language models have shown strong potential for video anomaly detection (VAD), yet the origin of their discriminative power remains unclear. We study this question using sparse mixture-of-experts (MoE) models, whose explicit expert structure and sparse activation make t …

نسخة أولية وصول مفتوح

Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs

Xiaojun Jia, Simeng Qin, Yiming Li وآخرون · 2026

Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples, especially in black-box settings where only open-source surrogate models are accessible. Existing targeted transfer attacks mainly align adversarial and target samples using global image-level features, such as encoder [CLS …

نسخة أولية وصول مفتوح

TempQ-Jail: Query-Constrained Candidate Ranking for Text-to-Video Jailbreak Attacks

Tianmeng Fang, Jiancheng Wang, Chen Wang وآخرون · 2026

Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video generation and security evaluation are costly, so an attacker often cannot test a large candidate pool. We therefore formulate T2V jailbreak as a query-constrained candidate …

المؤلفون المشاركون