الباحثون

Ke Zeng

المنشورات 1

نسخة أولية وصول مفتوح

DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents

Hanyang Wang, Zeyuan Liu, Zhengyu Chen وآخرون · 2026

On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stale before being used, …

المؤلفون المشاركون