الباحثون

Shuo Han

المنشورات 1

نسخة أولية وصول مفتوح

Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

Yu Li, Guangfeng Cai, Long-Fei Li وآخرون · 2026

Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not ex …

المؤلفون المشاركون