الباحثون

Alexandre Drouin

المنشورات 2

نسخة أولية وصول مفتوح

AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks

Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement. Despite their flexibility, their reliability on long tasks spanning multiple applications remains unclear. A trajectory, composed of …

نسخة أولية وصول مفتوح

AdaptArena: Evaluating Test-Time Personalization of Web Agents

Dongchan Shin, Xing Han Lù, Jiaqi Deng وآخرون · 2026

Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In practice, users rarely provide explicit profiles, requiring agents to infer latent prefere …

المؤلفون المشاركون