الباحثون

Han Luo

المنشورات 3

نسخة أولية وصول مفتوح

CIPO: Counterfactual Imagination Policy Optimization for Adaptive Tool Granularity Selection

Yu Li, Yunlu Wan, Zijian Zhu وآخرون · 2026

Large language model (LLM) agents solve complex tasks through multi-step interactions with external tools. These interactions often contain recurring local tool sequences. Treating such sequences as composite "Skills" can shorten tool-use trajectories and reduce repeated low-level decisions. However, when atomic tools …

نسخة أولية وصول مفتوح

Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

Yu Li, Guangfeng Cai, Long-Fei Li وآخرون · 2026

Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not ex …

المؤلفون المشاركون