الباحثون

Hanyang Wang

المنشورات 2

نسخة أولية وصول مفتوح

AgentGarten: Code Worlds for Evolving Agents

Jiawei Chi, Shangchen Miao, Zhiyuan Shi وآخرون · 2026

Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving bot …

نسخة أولية وصول مفتوح

DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents

Hanyang Wang, Zeyuan Liu, Zhengyu Chen وآخرون · 2026

On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stale before being used, …

المؤلفون المشاركون