الباحثون

Weijie Ren

المنشورات 4

نسخة أولية وصول مفتوح

ReTaCo: Residual-Target Control for On-Policy Distillation

Zixiang Ni, Zhuo Hu, Renjie Cao وآخرون · 2026

On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it …

نسخة أولية وصول مفتوح

SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration

Weijie Ren, Yanwen Zhang, Hao Li وآخرون · 2026

Multi-agent collaboration lets large language models (LLMs) improve question answering through deliberation and feedback. Yet shared discussion couples correction with exposure to the same mistakes, which can erode the diversity needed for voting. Self-consistency offers sampling diversity without feedback, while singl …

نسخة أولية وصول مفتوح

OPSRD: On-Policy Self-Role Distillation

Weijie Ren, Yanwen Zhang, Hao Li وآخرون · 2026

Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transf …

نسخة أولية وصول مفتوح

The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation

Hao Li, Meijia Chen, Weijie Ren وآخرون · 2026

On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however …

المؤلفون المشاركون