الباحثون

Zhuoning Guo

المنشورات 1

نسخة أولية وصول مفتوح

GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents

Bohan Lin, Liyi Chen, Zhuoning Guo وآخرون · 2026

On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, o …

المؤلفون المشاركون