نسخة أولية وصول مفتوح
Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation
On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection across RL- and SFT-continuation pairs and cross-model settings. …