الباحثون

Donghan Li

المنشورات 2

نسخة أولية وصول مفتوح

Suppressing Pressure, Amplifying Evidence: Self-Guided Attention Steering to Mitigate Sycophancy and Stubbornness

Yinghao He, Mengyu Xu, Haixiang Sun وآخرون · 2026

Reliable language models should resist unsupported user pressure while effectively using objective contextual information. However, models may exhibit sycophancy by yielding to unsupported user pressure or contextual stubbornness by failing to update their answers when relevant contextual information warrants revision. …

نسخة أولية وصول مفتوح

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

Hao Li, Meijia Chen, Weijie Ren وآخرون · 2026

On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however …

المؤلفون المشاركون