نسخة أولية وصول مفتوح
AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement. Despite their flexibility, their reliability on long tasks spanning multiple applications remains unclear. A trajectory, composed of …