الباحثون

Zhengyu Fang

المنشورات 3

نسخة أولية وصول مفتوح

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Zhengyu Fang, Seoyeon Hong, Jie Yang وآخرون · 2026

On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialis …

نسخة أولية وصول مفتوح

PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning

Muyang Li, Jie Yang, Zhengyu Fang وآخرون · 2026

Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advanta …

نسخة أولية وصول مفتوح

LastOPD: Taming Collapse in Latent On-Policy Distillation

Jie Yang, Zhengyu Fang, Zelin Xu وآخرون · 2026

On-policy distillation (OPD) corrects a student on the responses it writes, but its signal is the teacher's next-token distribution: it tells the student what the teacher says but misses how it thinks. Latent supervision promises the missing part by aligning the student's latent states to the teacher's. Recent methods …

المؤلفون المشاركون