Preprint Open access
GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents
On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, o …