نسخة أولية وصول مفتوح
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: th …