الباحثون

Jiajun Chai

المنشورات 2

نسخة أولية وصول مفتوح

SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation

Miteto Wei, Xiaohan Wang, Zehao Chen وآخرون · 2026

On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-cons …

نسخة أولية وصول مفتوح

TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs

Zihan Lin, Xiaohan Wang, Jie Cao وآخرون · 2026

Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of ex …

المؤلفون المشاركون