الباحثون

Liyi Chen

المنشورات 2

نسخة أولية وصول مفتوح

GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents

Bohan Lin, Liyi Chen, Zhuoning Guo وآخرون · 2026

On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, o …

نسخة أولية وصول مفتوح

RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation

Zixuan Yang, Yiqun Chen, Qi Liu وآخرون · 2026

Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuff …

المؤلفون المشاركون