الباحثون

Shengtian Yang

المنشورات 3

نسخة أولية وصول مفتوح

Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

Yu Li, Guangfeng Cai, Long-Fei Li وآخرون · 2026

Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not ex …

نسخة أولية وصول مفتوح

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Yu Li, Zheng Zhang, Xin Liu وآخرون · 2026

Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more …

المؤلفون المشاركون