الباحثون

Lu Sun

المنشورات 3

نسخة أولية وصول مفتوح

Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance

Rui Wen, Jiayang Liu, Zeyu Yang وآخرون · 2026

When a prompt injection attack succeeds, a Large Language Model (LLM) abandons its assigned system role to comply with an adversarial instruction. While prior work has extensively quantified how often this occurs, we ask a more fundamental question: where inside the network does the model actually decide to break the r …

نسخة أولية وصول مفتوح

Near-Duplicate Families Break Exact-Record Membership Inference

Yiyong Liu, Jiayang Liu, Yixin Tan وآخرون · 2026

Membership inference (MI) asks whether a specific record appeared in a model's training set and is increasingly used as evidence for data provenance and copyright auditing. These applications require determining whether the exact queried record was used for training, rather than merely whether the model was exposed to …

المؤلفون المشاركون