Abstract
Vertical-domain few-shot classification remains challenging for small language models, as limited supervision makes it difficult to acquire domain-specific decision knowledge. On-Policy Distillation (OPD) can improve teacher-guided adaptation by supervising student-generated rollouts, while GRPO-based reinforcement learning can further refine downstream predictions. However, existing KD-to-RL pipelines typically rely on globally fixed transition schedules, ignoring that different samples may require different amounts of teacher-guided acquisition before reward-driven refinement. We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity. PIVOT moves low-perplexity samples to GRPO for reward-driven refinement while keeping high-perplexity samples under OPD for continued domain knowledge acquisition. Experiments on Banking77 and HWU64 show that PIVOT consistently outperforms continued OPD and globally synchronized OPD$\rightarrow$GRPO baselines under the same number of post-warm-up student optimization steps, achieving stronger downstream performance and more stable training dynamics.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Li, H., Zhang, Y., Cheng, N., Li, Z., Zhu, Y., Wang, Y., Wang, S., & Xiao, J. (2026). PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation. https://omanscience.com/en/articles/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few-shot-distillation
MLA 9
Li, Heng, et al. "PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation." https://omanscience.com/en/articles/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few-shot-distillation.
Chicago (author–date)
Li, Heng, Yong Zhang, Ning Cheng, Zhigen Li, Yun Zhu, Yanmeng Wang, Shaojun Wang, and Jing Xiao. 2026. "PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation." https://omanscience.com/en/articles/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few-shot-distillation.
Harvard
Li, H., Zhang, Y., Cheng, N., Li, Z., Zhu, Y., Wang, Y., Wang, S. and Xiao, J. (2026) 'PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation', Available at: https://omanscience.com/en/articles/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few-shot-distillation.
Vancouver
Li H, Zhang Y, Cheng N, Li Z, Zhu Y, Wang Y, et al. PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation. https://omanscience.com/en/articles/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few-shot-distillation
IEEE
H. Li, Y. Zhang, N. Cheng, Z. Li, Y. Zhu, Y. Wang, S. Wang, and J. Xiao, "PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation," https://omanscience.com/en/articles/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few-shot-distillation.