Abstract

Activation steering offers an inference-time defense for vision--language models (VLMs) by modifying intermediate representations without updating backbone parameters. However, protection on benchmark inputs may not persist across alternative expressions of the same harmful request. We investigate this gap using fixed textual, visual, and joint reformulations designed to preserve the underlying intent, and find that these changes can bypass representative steering defenses. A complementary local analysis provides a sufficient condition under which a reformulation can cross a surrogate safety margin despite any admissible change in the local steering correction. We then introduce SteerProbe, an output-only black-box attack that learns to select effective reformulations for unseen requests from a shared calibration budget. Across three VLM backbones, two benchmarks, and three steering defenses, SteerProbe increases Harmful Rate in all defended settings using 500 total calibration queries per endpoint and benchmark, raising the average from 7.43% to 18.36%. These findings highlight that robustness on original benchmark inputs is insufficient to characterize the safety of steering defenses and motivate reformulation robustness as an important evaluation dimension. They further motivate steering mechanisms that preserve safety across intent-preserving multimodal variations while maintaining benign utility.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, X., Hu, A., Liu, H., Pang, S., Ye, Q., & Hu, H. (2026). SteerProbe: Learning to Bypass Safety Steering in Vision-Language Models. https://omanscience.com/en/articles/steerprobe-learning-to-bypass-safety-steering-in-vision-language-models

MLA 9

Zhang, Xinwei, et al. "SteerProbe: Learning to Bypass Safety Steering in Vision-Language Models." https://omanscience.com/en/articles/steerprobe-learning-to-bypass-safety-steering-in-vision-language-models.

Chicago (author–date)

Zhang, Xinwei, Aoting Hu, Hangcheng Liu, Shuchao Pang, Qingqing Ye, and Haibo Hu. 2026. "SteerProbe: Learning to Bypass Safety Steering in Vision-Language Models." https://omanscience.com/en/articles/steerprobe-learning-to-bypass-safety-steering-in-vision-language-models.

Harvard

Zhang, X., Hu, A., Liu, H., Pang, S., Ye, Q. and Hu, H. (2026) 'SteerProbe: Learning to Bypass Safety Steering in Vision-Language Models', Available at: https://omanscience.com/en/articles/steerprobe-learning-to-bypass-safety-steering-in-vision-language-models.

Vancouver

Zhang X, Hu A, Liu H, Pang S, Ye Q, Hu H. SteerProbe: Learning to Bypass Safety Steering in Vision-Language Models. https://omanscience.com/en/articles/steerprobe-learning-to-bypass-safety-steering-in-vision-language-models

IEEE

X. Zhang, A. Hu, H. Liu, S. Pang, Q. Ye, and H. Hu, "SteerProbe: Learning to Bypass Safety Steering in Vision-Language Models," https://omanscience.com/en/articles/steerprobe-learning-to-bypass-safety-steering-in-vision-language-models.