Preprint Open access
USA: Update-aware SAM for Cross-domain On-Policy Disitllation of Language Agents
On-policy distillation instils multi-turn agentic reasoning through dense token-level supervision on the student's own trajectories, but a single domain saturates early, so further supervision has to be drawn from other domains. Multi-domain data mixing is the most direct way of incorporating them, at the cost of confl …