Authors

Weijie Ren

Publications 4

Preprint Open access

ReTaCo: Residual-Target Control for On-Policy Distillation

Zixiang Ni, Zhuo Hu, Renjie Cao et al. · 2026

On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it …

Preprint Open access

OPSRD: On-Policy Self-Role Distillation

Weijie Ren, Yanwen Zhang, Hao Li et al. · 2026

Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transf …

Preprint Open access

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

Hao Li, Meijia Chen, Weijie Ren et al. · 2026

On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however …

Co-authors