نسخة أولية وصول مفتوح
ResOPD: Tail Residualization for Sparse On-Policy Distillation
On-policy distillation (OPD) trains a student model on self-generated trajectories, but transmitting dense teacher distributions across long reasoning traces creates prohibitive communication and memory bottlenecks. Practical systems therefore rely on sparse teacher interfaces, typically transmitting either the sampled …