الباحثون

Pengjun Xie

المنشورات 3

نسخة أولية وصول مفتوح

Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents

Yingjian Zhu, Zhenyi Wang, Jiaxin Guo وآخرون · 2026

Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on ground-truth answer …

نسخة أولية وصول مفتوح

ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

Qiang Zhang, Ruixue Ding, Fanrui Zhang وآخرون · 2026

Reinforcement learning has substantially improved large language model (LLM) agents in verifiable domains, but remains difficult to apply to open-ended agent tasks, where solutions are diverse and reliable scalar rewards are hard to obtain. Recent pairwise evaluation methods alleviate reward discrimination collapse by …

المؤلفون المشاركون