Abstract
A vision-language model may need to combine an image with a prompt to recognize a safety risk that neither reveals alone. Where does this joint safety judgment become accessible inside the model? We introduce SSU-Bench, a dataset of matched safe and unsafe image-text combinations constructed using single-item prompt edits or image edits with annotated intended regions. Using three vision-language models, we transfer internal states between paired inputs and measure the resulting change in the safety verdict. Across models and both types of counterfactual, interventions at the changed input positions are effective in earlier decoder layers, while interventions at the final input token become effective later. Directions estimated from other examples produce similar late-layer effects. A linear readout of the final-token state also predicts the model's own verdict, including incorrect judgments, and cross-model comparisons reveal similarities in the patterns of counterfactual change. These findings identify a recurring transition in where interventions can influence a joint safety verdict and distinguish a readable model decision from a correct safety judgment.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Tarzjani, F. D., Wijewardena, M., Romanus, A., & Mohanty, S. (2026). From Local Evidence to Safety Verdicts: Causal Tracing in Vision-Language Models. https://omanscience.com/en/articles/from-local-evidence-to-safety-verdicts-causal-tracing-in-vision-language-models
MLA 9
Tarzjani, Faezeh Dehghan, et al. "From Local Evidence to Safety Verdicts: Causal Tracing in Vision-Language Models." https://omanscience.com/en/articles/from-local-evidence-to-safety-verdicts-causal-tracing-in-vision-language-models.
Chicago (author–date)
Tarzjani, Faezeh Dehghan, Mevan Wijewardena, Alexander Romanus, and Sampad Mohanty. 2026. "From Local Evidence to Safety Verdicts: Causal Tracing in Vision-Language Models." https://omanscience.com/en/articles/from-local-evidence-to-safety-verdicts-causal-tracing-in-vision-language-models.
Harvard
Tarzjani, F. D., Wijewardena, M., Romanus, A. and Mohanty, S. (2026) 'From Local Evidence to Safety Verdicts: Causal Tracing in Vision-Language Models', Available at: https://omanscience.com/en/articles/from-local-evidence-to-safety-verdicts-causal-tracing-in-vision-language-models.
Vancouver
Tarzjani FD, Wijewardena M, Romanus A, Mohanty S. From Local Evidence to Safety Verdicts: Causal Tracing in Vision-Language Models. https://omanscience.com/en/articles/from-local-evidence-to-safety-verdicts-causal-tracing-in-vision-language-models
IEEE
F. D. Tarzjani, M. Wijewardena, A. Romanus, and S. Mohanty, "From Local Evidence to Safety Verdicts: Causal Tracing in Vision-Language Models," https://omanscience.com/en/articles/from-local-evidence-to-safety-verdicts-causal-tracing-in-vision-language-models.