Authors

Shayan Mohajer Hamidi

Publications 3

Preprint Open access

Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts

Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both s …

Preprint Open access

Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents

Hierarchical reinforcement learning improves long-horizon control by organizing primitive actions around persistent subgoals and assigning credit at multiple temporal scales. Recent hierarchical language agents bring these benefits to interactive tasks by explicitly separating subgoal planning from action execution. We …

Preprint Open access

Smaller Models, Better Rejects: Preference Distillation Scaling

Rui Cai, Wenhui Zhu, Xiwen Chen et al. · 2026

Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither a …

Co-authors