الباحثون

Qizhou Wang

المنشورات 8

نسخة أولية وصول مفتوح

Benchmarking Behavioral Steerability in Behavior Foundation Models

Minghe Gao, Zhanxi Yan, Jiahui Liu وآخرون · 2026

Behavior Foundation Models (BFMs) are emerging as a paradigm for translating human intentions into executable humanoid behaviors. As these models evolve beyond behavior generation toward general-purpose behavioral systems, a fundamental question arises: can they be reliably steered according to user intentions? In this …

نسخة أولية وصول مفتوح

TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

Fengpeng Li, Kemou Li, Qizhou Wang وآخرون · 2026

Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which …

نسخة أولية وصول مفتوح

Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing

Kemou Li, Qizhou Wang, Yue Wang وآخرون · 2026

Open-weight LLMs are released not only as fixed products but also as substrates for downstream fine-tuning. This openness, however, creates legal and ethical risks because users may misuse fine-tuning to instill illicit knowledge or enable hostile operations. Model providers therefore need apre-release defense against …

نسخة أولية وصول مفتوح

LLM Persona Unlearning

Kemou Li, Zhuan Shi, Qizhou Wang وآخرون · 2026

Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-training teaches conditional enactment and makes a helpful Assistant the default, but it does not erase alternative modes from the weights; explicit prompts can therefor …

نسخة أولية وصول مفتوح

UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?

Puning Yang, Qizhou Wang, Junchi Yu وآخرون · 2026

Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this issue, existing unlearning methods typically rely on training-based parameter updates, such as gradient ascent and its variants, to delete targeted content while preservin …

نسخة أولية وصول مفتوح

VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation

Hanxun Huang, Yutao Wu, Qizhou Wang وآخرون · 2026

Large language models (LLMs) have made misinformation inexpensive to produce but not to verify, creating a growing asymmetry in the information ecosystem. Under tight time, labor, and budget constraints, media organizations, platforms, and fact-checkers rely on screening to prioritize which content to verify. We introd …

المؤلفون المشاركون