Abstract

Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In practice, users rarely provide explicit profiles, requiring agents to infer latent preferences from implicit signals. Despite its importance for deployment, this problem setting is largely underexplored in existing benchmarks. To address this gap, we introduce AdaptArena, a benchmark for evaluating test-time personalization of web agents via implicit preference inference. AdaptArena consists of 480 tasks, featuring both single-preference and double-preference scenarios. Each evaluation task must be solved by retrieving and leveraging the most relevant historical user trajectory that implicitly encodes the target preference. In addition, we introduce AdaptiveAgent, a retrieval-based framework for standardized evaluation of implicit preference inference. Experiments reveal a substantial performance gap: while oracle agents with access to ground-truth preferences achieve an 82.92% success rate, the evaluated LLM agents using our framework reach at most 15.62%. Furthermore, we find that correctly inferring user preferences is necessary but not sufficient for task success, as execution failures in downstream web interactions remain a significant bottleneck even when agents align with the target preference. These findings highlight implicit preference inference and robust action grounding as key challenges for deploying reliable, user-facing web agents. We release our code: https://github.com/McGill-NLP/web-agents-test-time-adaptations

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Shin, D., Lù, X. H., Deng, J., Gala, J., Browne, T. V., Moon, J., Liu, F., Drouin, A., Reddy, S., & Lacoste, A. (2026). AdaptArena: Evaluating Test-Time Personalization of Web Agents. https://omanscience.com/en/articles/adaptarena-evaluating-test-time-personalization-of-web-agents

MLA 9

Shin, Dongchan, et al. "AdaptArena: Evaluating Test-Time Personalization of Web Agents." https://omanscience.com/en/articles/adaptarena-evaluating-test-time-personalization-of-web-agents.

Chicago (author–date)

Shin, Dongchan, Xing Han Lù, Jiaqi Deng, Jay Gala, Tomás Vergara Browne, Jaewon Moon, Fengyuan Liu, Alexandre Drouin, Siva Reddy, and Alexandre Lacoste. 2026. "AdaptArena: Evaluating Test-Time Personalization of Web Agents." https://omanscience.com/en/articles/adaptarena-evaluating-test-time-personalization-of-web-agents.

Harvard

Shin, D., Lù, X. H., Deng, J., Gala, J., Browne, T. V., Moon, J., Liu, F., Drouin, A., Reddy, S. and Lacoste, A. (2026) 'AdaptArena: Evaluating Test-Time Personalization of Web Agents', Available at: https://omanscience.com/en/articles/adaptarena-evaluating-test-time-personalization-of-web-agents.

Vancouver

Shin D, Lù XH, Deng J, Gala J, Browne TV, Moon J, et al. AdaptArena: Evaluating Test-Time Personalization of Web Agents. https://omanscience.com/en/articles/adaptarena-evaluating-test-time-personalization-of-web-agents

IEEE

D. Shin, X. H. Lù, J. Deng, J. Gala, T. V. Browne, J. Moon, F. Liu, A. Drouin, S. Reddy, and A. Lacoste, "AdaptArena: Evaluating Test-Time Personalization of Web Agents," https://omanscience.com/en/articles/adaptarena-evaluating-test-time-personalization-of-web-agents.