[
    {
        "id": "osp-19632",
        "type": "article-journal",
        "title": "Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation",
        "author": [
            {
                "family": "Chen",
                "given": "Xiang"
            },
            {
                "family": "Su",
                "given": "Futao"
            },
            {
                "family": "Wang",
                "given": "Kong"
            },
            {
                "family": "Chen",
                "given": "Jiayi"
            },
            {
                "family": "Li",
                "given": "TanLin"
            }
        ],
        "URL": "https://omanscience.com/en/articles/spend-teacher-tokens-where-they-matter-success-referenced-on-policy-distillation",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher, but providing such supervision for every rollout requires substantial teacher computation. We introduce Success-Referenced On-Policy Distillation (SR-OPD), which reduces this cost by selecting which prompts and rollouts receive teacher supervision. When the student produces both successful and failed rollouts for the same prompt, a successful rollout can serve as a natural reference for selecting failed rollouts. SR-OPD therefore focuses on such prompts and prioritizes failed rollouts whose hidden-state trajectories show sustained divergence from a successful reference, while accounting for estimated teacher-input cost. Across three teacher-student pairs and six mathematical reasoning benchmarks, SR-OPD uses only 3.46-5.02% of the teacher-input tokens required by Vanilla OPD in the one-pass setting while maintaining comparable reasoning performance. Under a controlled setting matched to 5% of Vanilla OPD's teacher-input budget, further experiments support both key design choices: focusing supervision on prompts with both successful and failed rollouts, and using successful rollouts to guide failure selection. These results indicate that a student's own successful behavior can serve as a useful reference for allocating teacher supervision under a fixed teacher-input budget."
    }
]