Preprint Open access
ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
Large language model agents are increasingly deployed to perform complex tasks in real-world environments. However, the knowledge required for correct behavior in these environments is often implicit, undisclosed, and subject to change over time. Recent continual-learning harnesses seek to address this challenge by ena …