الباحثون

Xuebo Liu

المنشورات 2

نسخة أولية وصول مفتوح

SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction

Hexuan Deng, Yue Wang, Wenyu Jiang وآخرون · 2026

Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repe …

نسخة أولية وصول مفتوح

GRPODropout: Less is More for Online Reinforcement Learning Rollouts

Hexuan Deng, Zihao Yan, Xuebo Liu وآخرون · 2026

Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such a …

المؤلفون المشاركون