الباحثون

Daksh Raghuvanshi

المنشورات 1

نسخة أولية وصول مفتوح

StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents

Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live- …

المؤلفون المشاركون