الباحثون

Zhanyang Jin

المنشورات 2

نسخة أولية وصول مفتوح

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Siqi Zhu, Suozhi Huang, Kaixuan Zhang وآخرون · 2026

Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, op …

نسخة أولية وصول مفتوح

Learning from Teacher Continuations at Student States

Haojin Wang, Dylan Zhang, Huaibo Chen وآخرون · 2026

We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillatio …

المؤلفون المشاركون