الباحثون

Fengpeng Li

المنشورات 4

نسخة أولية وصول مفتوح

TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

Fengpeng Li, Kemou Li, Qizhou Wang وآخرون · 2026

Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which …

نسخة أولية وصول مفتوح

Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing

Kemou Li, Qizhou Wang, Yue Wang وآخرون · 2026

Open-weight LLMs are released not only as fixed products but also as substrates for downstream fine-tuning. This openness, however, creates legal and ethical risks because users may misuse fine-tuning to instill illicit knowledge or enable hostile operations. Model providers therefore need apre-release defense against …

نسخة أولية وصول مفتوح

LLM Persona Unlearning

Kemou Li, Zhuan Shi, Qizhou Wang وآخرون · 2026

Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-training teaches conditional enactment and makes a helpful Assistant the default, but it does not erase alternative modes from the weights; explicit prompts can therefor …

المؤلفون المشاركون