[
    {
        "id": "osp-17433",
        "type": "article-journal",
        "title": "When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better",
        "author": [
            {
                "family": "Zhao",
                "given": "Siyan"
            },
            {
                "family": "Fu",
                "given": "Yonggan"
            },
            {
                "family": "Jiang",
                "given": "Jindong"
            },
            {
                "family": "Liu",
                "given": "Shih-Yang"
            },
            {
                "family": "Bian",
                "given": "Song"
            },
            {
                "family": "Lee",
                "given": "Byung-Kwan"
            },
            {
                "family": "Sreenivas",
                "given": "Sharath Turuvekere"
            },
            {
                "family": "Dai",
                "given": "Wenliang"
            },
            {
                "family": "Ye",
                "given": "Hanrong"
            },
            {
                "family": "Grover",
                "given": "Aditya"
            },
            {
                "family": "Molchanov",
                "given": "Pavlo"
            }
        ],
        "URL": "https://omanscience.com/en/articles/when-do-we-need-on-policy-distillation-distilling-on-offline-student-rollouts-is-often-better",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models. In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs? We show that a simple alternative, Semi-OPD, which distills from offline rollouts generated by the initial student, can often outperform OPD in both accuracy and training efficiency. Across 17 teacher-student pairs ranging from 1.5B to 235B parameters, Semi-OPD outperforms OPD in 14 cases, with up to +13.6% accuracy and 11.4x training speedup. We further find that the choice between OPD and Semi-OPD depends on the alignment between the initial teacher and student, quantified by an output-token overlap ratio: OPD is beneficial only when the two are highly aligned with high overlap ratios. Our deeper investigation suggests that effective distillation requires on-policyness w.r.t. both the student and the teacher. For misaligned pairs, student rollouts can become increasingly off-policy w.r.t. the teacher as context length grows, weakening the distillation signal. In contrast, Semi-OPD is often more stable, as it distills on shorter contexts while covering full trajectories and exposing the student to more teacher-preferred tokens. Beyond proposing Semi-OPD as an efficient alternative, our work motivates the community to rethink when to use OPD and to study stronger OPD variants with meaningful teacher-student pairs."
    }
]