Preprint Open access
PG-SFT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning
Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., general reasoning, tool calling, code generation) of the base model. In this work, we focu …