Abstract
Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not explicitly trace the read-write dependencies through which commands affect the final outcome. Consequently, training signals could still be assigned to irrelevant operations, weakening learning from relevant steps. In this paper, we propose Dependency-Aware Group Policy Optimization (DepGPO), which uses execution dependencies between commands to guide credit assignment for terminal agents. Specifically, we construct a command dependency graph from execution traces and trace backward from the resources inspected by the task verifier. We then assign credit to relevant writes and their supporting reads along these paths, and use it to redistribute trajectory advantages across steps. Extensive comparative experiments and ablation studies demonstrate that DepGPO improves task performance and training stability on complex terminal tasks.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Li, Y., Cai, G., Li, L. F., Han, S., Yang, S., Luo, H., Yang, K., & Feng, L. (2026). Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents. https://omanscience.com/en/articles/credit-where-it-matters-dependency-aware-policy-optimization-for-terminal-agents
MLA 9
Li, Yu, et al. "Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents." https://omanscience.com/en/articles/credit-where-it-matters-dependency-aware-policy-optimization-for-terminal-agents.
Chicago (author–date)
Li, Yu, Guangfeng Cai, Long-Fei Li, Shuo Han, Shengtian Yang, Han Luo, Kaibing Yang, and Lei Feng. 2026. "Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents." https://omanscience.com/en/articles/credit-where-it-matters-dependency-aware-policy-optimization-for-terminal-agents.
Harvard
Li, Y., Cai, G., Li, L. F., Han, S., Yang, S., Luo, H., Yang, K. and Feng, L. (2026) 'Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents', Available at: https://omanscience.com/en/articles/credit-where-it-matters-dependency-aware-policy-optimization-for-terminal-agents.
Vancouver
Li Y, Cai G, Li LF, Han S, Yang S, Luo H, et al. Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents. https://omanscience.com/en/articles/credit-where-it-matters-dependency-aware-policy-optimization-for-terminal-agents
IEEE
Y. Li, G. Cai, L. F. Li, S. Han, S. Yang, H. Luo, K. Yang, and L. Feng, "Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents," https://omanscience.com/en/articles/credit-where-it-matters-dependency-aware-policy-optimization-for-terminal-agents.