الملخص
Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions. Existing visual red-teaming approaches use adversarial visual content to manipulate this process. However, they primarily target model inference and do not explicitly account for structured input processing or action post-processing. Consequently, model-level success does not establish control over browser execution and cannot reliably characterize end-to-end agent robustness. To address this gap, we formulate red teaming for vision-grounded web agents as an end-to-end grounding-to-execution problem, and introduce WebMirage, a framework that crafts localized visual perturbations that cause agents to select attacker-controlled content and execute the corresponding browser action across varying webpage renderings. It uses a role-slot abstraction and webpage recomposition to capture competition among webpage elements, and dataflow analysis to align optimization with action post-processing. We evaluate WebMirage across four agent configurations and six VLM backbones on 2,250 tasks covering 13 public websites and a sandbox benchmark. WebMirage achieves an average attack success rate of 91.9%, compared with 17.4% for the strongest baseline, and remains effective against three agent-level defenses.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Han, W., Li, L. T., Zhang, M., Jiang, Y., & Tao, G. (2026). Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution. https://omanscience.com/ar/articles/adversarial-images-hijack-web-agents-from-visual-grounding-to-browser-execution
MLA 9
Han, Wanjing, et al. "Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution." https://omanscience.com/ar/articles/adversarial-images-hijack-web-agents-from-visual-grounding-to-browser-execution.
شيكاغو (المؤلف–التاريخ)
Han, Wanjing, Levi Taiji Li, Mu Zhang, Yue Jiang, and Guanhong Tao. 2026. "Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution." https://omanscience.com/ar/articles/adversarial-images-hijack-web-agents-from-visual-grounding-to-browser-execution.
هارفارد
Han, W., Li, L. T., Zhang, M., Jiang, Y. and Tao, G. (2026) 'Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution', Available at: https://omanscience.com/ar/articles/adversarial-images-hijack-web-agents-from-visual-grounding-to-browser-execution.
فانكوفر
Han W, Li LT, Zhang M, Jiang Y, Tao G. Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution. https://omanscience.com/ar/articles/adversarial-images-hijack-web-agents-from-visual-grounding-to-browser-execution
IEEE
W. Han, L. T. Li, M. Zhang, Y. Jiang, and G. Tao, "Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution," https://omanscience.com/ar/articles/adversarial-images-hijack-web-agents-from-visual-grounding-to-browser-execution.