الباحثون

Ali Holmov

المنشورات 1

نسخة أولية وصول مفتوح

User Model Extraction via Belief Self-Distillation

Ali Holmov, Yiran Huang, Kirill Bykov وآخرون · 2026

Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact us …

المؤلفون المشاركون