نسخة أولية وصول مفتوح
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed it …