الباحثون

Jun Sakuma

المنشورات 3

نسخة أولية وصول مفتوح

Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance

Rui Wen, Jiayang Liu, Zeyu Yang وآخرون · 2026

When a prompt injection attack succeeds, a Large Language Model (LLM) abandons its assigned system role to comply with an adversarial instruction. While prior work has extensively quantified how often this occurs, we ask a more fundamental question: where inside the network does the model actually decide to break the r …

نسخة أولية وصول مفتوح

No Free Efficiency: Revisiting the Trade-off Between Training Efficiency and Model Vulnerability

Yiyong Liu, Jun Sakuma, Michael Backes وآخرون · 2026

Training efficiency has become the central driver of recent progress in foundation models. To overcome the massive computational and data requirements of large-scale training, researchers increasingly adopt strategies such as selective data sampling, efficient pre-training, and simplified reinforcement learning pipelin …

المؤلفون المشاركون