الباحثون

Xiangyu Peng

المنشورات 2

نسخة أولية وصول مفتوح

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

Jixuan Chen, Jiaxin Zhang, Qinyuan Ye وآخرون · 2026

Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This …

نسخة أولية وصول مفتوح

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

Jiaxin Zhang, Xiangyu Peng, Qinglin Chen وآخرون · 2026

Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating training principle that aug …

المؤلفون المشاركون