Preprint Open access
Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoders, which learn one f …