Authors

Zhengyu Fang

Publications 3

Preprint Open access

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialis …

Preprint Open access

LastOPD: Taming Collapse in Latent On-Policy Distillation

Jie Yang, Zhengyu Fang, Zelin Xu et al. · 2026

On-policy distillation (OPD) corrects a student on the responses it writes, but its signal is the teacher's next-token distribution: it tells the student what the teacher says but misses how it thinks. Latent supervision promises the missing part by aligning the student's latent states to the teacher's. Recent methods …

Co-authors