Abstract

Small language models are often post-trained as students on reasoning traces from stronger teacher models to efficiently learn new skills. However, token-level imitation on traces that lie far outside the student's expected distribution often produces \textit{confident conflicts}, whereby the student is required to imitate a continuation that it deems unlikely (i.e., low-probability) despite being confident in a different continuation (i.e., in a low-entropy state). To mitigate the degradation in generalisation and catastrophic forgetting caused by these conflicts, we propose \textbf{Entropy-Aware Mixing}: a dynamic per-token interpolation of the student and teacher distributions, gated by the student's predictive entropy. We implement both convex and geometric interpolations for both offline trace generation (via speculative decoding, then SFT) and on-policy forward-KL distillation. Our results show that entropy-aware mixing stabilises distillation, improving in-distribution and out-of-distribution math reasoning while better preserving general capabilities than fixed-teacher supervision. Nonetheless, the optimal entropy schedule depends on the training source, with offline-generated traces favouring concave schedules (greater overall teacher influence) and on-policy training favouring linear or convex schedules (teacher concentrated in high-entropy states).

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Giraldo, J. G., Santelmo, M., Durech, E., Schlag, I., Pyatkin, V., & Bosselut, A. (2026). Learning What to Imitate: Entropy-Aware Distribution Mixing. https://omanscience.com/en/articles/learning-what-to-imitate-entropy-aware-distribution-mixing

MLA 9

Giraldo, Juan Garcia, et al. "Learning What to Imitate: Entropy-Aware Distribution Mixing." https://omanscience.com/en/articles/learning-what-to-imitate-entropy-aware-distribution-mixing.

Chicago (author–date)

Giraldo, Juan Garcia, Matteo Santelmo, Eduard Durech, Imanol Schlag, Valentina Pyatkin, and Antoine Bosselut. 2026. "Learning What to Imitate: Entropy-Aware Distribution Mixing." https://omanscience.com/en/articles/learning-what-to-imitate-entropy-aware-distribution-mixing.

Harvard

Giraldo, J. G., Santelmo, M., Durech, E., Schlag, I., Pyatkin, V. and Bosselut, A. (2026) 'Learning What to Imitate: Entropy-Aware Distribution Mixing', Available at: https://omanscience.com/en/articles/learning-what-to-imitate-entropy-aware-distribution-mixing.

Vancouver

Giraldo JG, Santelmo M, Durech E, Schlag I, Pyatkin V, Bosselut A. Learning What to Imitate: Entropy-Aware Distribution Mixing. https://omanscience.com/en/articles/learning-what-to-imitate-entropy-aware-distribution-mixing

IEEE

J. G. Giraldo, M. Santelmo, E. Durech, I. Schlag, V. Pyatkin, and A. Bosselut, "Learning What to Imitate: Entropy-Aware Distribution Mixing," https://omanscience.com/en/articles/learning-what-to-imitate-entropy-aware-distribution-mixing.