الباحثون

Purui Liu

المنشورات 1

نسخة أولية وصول مفتوح

RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation

Hengbo Xiao, Boyao Zhang, Purui Liu وآخرون · 2026

Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in domains where task outcomes can be reliably evaluated, but long-horizon interaction remains challenging due to sparse terminal feedback and difficult credit assignment. Process rewards provide denser supervision, yet the capabiliti …

المؤلفون المشاركون