Abstract
Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives of their own? Online marketplaces, for example, may favor some products over others, steering agents away from the user's objective. Existing CUA benchmarks cover cooperative settings or explicit attacks, but do not test whether agents preserve user objectives when the environment itself has a stake in the outcome. We introduce CAVEAT, a controlled benchmark spanning nine marketplace environments and a taxonomy of eight common steering mechanisms. Across five model families, agents purchase the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering mechanisms are enabled. Larger models and more reasoning improve robustness, but substantial failures persist. Trajectory analysis and targeted ablations identify three weaknesses in how agents decide: they (1) prematurely narrow the set of alternatives they consider, (2) impose priorities the user never stated, and (3) commit before resolving decision-relevant evidence. Guided by this diagnosis, we develop CAVEAT-Harness, which targets these failures and raises the optimal purchase rate by up to 80.0 percentage points, and show that targeted post-training further improves a smaller open model. These results establish incentive robustness as a distinct challenge for delegated agents, diagnose failure modes, and show how targeted interventions can substantially improve robustness.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Li, Y., Epperson, W., Deng, W., & Huang, Z. (2026). CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments. https://omanscience.com/en/articles/caveat-towards-robust-computer-use-agents-in-incentive-misaligned-environments
MLA 9
Li, Yuxuan, et al. "CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments." https://omanscience.com/en/articles/caveat-towards-robust-computer-use-agents-in-incentive-misaligned-environments.
Chicago (author–date)
Li, Yuxuan, Will Epperson, Wesley Deng, and Zezhou Huang. 2026. "CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments." https://omanscience.com/en/articles/caveat-towards-robust-computer-use-agents-in-incentive-misaligned-environments.
Harvard
Li, Y., Epperson, W., Deng, W. and Huang, Z. (2026) 'CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments', Available at: https://omanscience.com/en/articles/caveat-towards-robust-computer-use-agents-in-incentive-misaligned-environments.
Vancouver
Li Y, Epperson W, Deng W, Huang Z. CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments. https://omanscience.com/en/articles/caveat-towards-robust-computer-use-agents-in-incentive-misaligned-environments
IEEE
Y. Li, W. Epperson, W. Deng, and Z. Huang, "CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments," https://omanscience.com/en/articles/caveat-towards-robust-computer-use-agents-in-incentive-misaligned-environments.