الباحثون

Zhengze Zhou

المنشورات 4

نسخة أولية وصول مفتوح

Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both s …

نسخة أولية وصول مفتوح

Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents

Shayan Mohajer Hamidi, Yize Cheng, Yuanda Xu وآخرون · 2026

Hierarchical reinforcement learning improves long-horizon control by organizing primitive actions around persistent subgoals and assigning credit at multiple temporal scales. Recent hierarchical language agents bring these benefits to interactive tasks by explicitly separating subgoal planning from action execution. We …

نسخة أولية وصول مفتوح

Smaller Models, Better Rejects: Preference Distillation Scaling

Rui Cai, Wenhui Zhu, Xiwen Chen وآخرون · 2026

Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither a …

المؤلفون المشاركون