نسخة أولية وصول مفتوح
CIPO: Counterfactual Imagination Policy Optimization for Adaptive Tool Granularity Selection
Large language model (LLM) agents solve complex tasks through multi-step interactions with external tools. These interactions often contain recurring local tool sequences. Treating such sequences as composite "Skills" can shorten tool-use trajectories and reduce repeated low-level decisions. However, when atomic tools …