الباحثون

Qi Gu

المنشورات 2

نسخة أولية وصول مفتوح

COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning

Zicheng Hu, Zhijian Zhou, Xuan Zhang وآخرون · 2026

Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective. We show that this \emph{policy-side correction} alo …

نسخة أولية وصول مفتوح

Consistent Plan-Act for Long-Horizon Agentic Tasks

Heng-Zhuang Li, Yi-Kai Zhang, Yu Wang وآخرون · 2026

Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner and actor roles. To investigate coordination failures in these tasks, we prompt both agen …

المؤلفون المشاركون