الباحثون

Jiayang Liu

المنشورات 4

نسخة أولية وصول مفتوح

Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance

Rui Wen, Jiayang Liu, Zeyu Yang وآخرون · 2026

When a prompt injection attack succeeds, a Large Language Model (LLM) abandons its assigned system role to comply with an adversarial instruction. While prior work has extensively quantified how often this occurs, we ask a more fundamental question: where inside the network does the model actually decide to break the r …

نسخة أولية وصول مفتوح

Near-Duplicate Families Break Exact-Record Membership Inference

Yiyong Liu, Jiayang Liu, Yixin Tan وآخرون · 2026

Membership inference (MI) asks whether a specific record appeared in a model's training set and is increasingly used as evidence for data provenance and copyright auditing. These applications require determining whether the exact queried record was used for training, rather than merely whether the model was exposed to …

نسخة أولية وصول مفتوح

TempQ-Jail: Query-Constrained Candidate Ranking for Text-to-Video Jailbreak Attacks

Tianmeng Fang, Jiancheng Wang, Chen Wang وآخرون · 2026

Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video generation and security evaluation are costly, so an attacker often cannot test a large candidate pool. We therefore formulate T2V jailbreak as a query-constrained candidate …

المؤلفون المشاركون