Abstract

How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR). Across four models, we identify attention heads causally necessary for OCR, and discover that these are in fact general-purpose heads that output interpretable semantic features across all image tokens. For example, pointing these heads at an image token containing the word "bike" causes Qwen3-VL-8B to output "bike," but pointing them at a bird wing causes the model to output the token "feathers." We collapse these heads' attention weights into a single verbalization lens transformation that reveals interpretable semantic features in hidden states across all layers. When combined with projection to vocabulary space, we can obtain interpretable labels starting from layer 0, showing that image representations are in fact aligned with language in early layers. We find that we can also use the inverse of this transformation to edit non-word concepts, e.g., replacing a tractor with a revolver in a naturalistic image, providing causal evidence that this subspace is useful for more than just OCR. Our results are an example of how the study of specific mechanisms can shed light on broader interpretability problems.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Feucht, S., Krojer, B., Wang, S., Abrahamsen, H., Wallace, B. C., & Bau, D. (2026). Using OCR Heads to Verbalize Image Semantics. https://omanscience.com/en/articles/using-ocr-heads-to-verbalize-image-semantics

MLA 9

Feucht, Sheridan, et al. "Using OCR Heads to Verbalize Image Semantics." https://omanscience.com/en/articles/using-ocr-heads-to-verbalize-image-semantics.

Chicago (author–date)

Feucht, Sheridan, Benno Krojer, Sarah Wang, Henry Abrahamsen, Byron C. Wallace, and David Bau. 2026. "Using OCR Heads to Verbalize Image Semantics." https://omanscience.com/en/articles/using-ocr-heads-to-verbalize-image-semantics.

Harvard

Feucht, S., Krojer, B., Wang, S., Abrahamsen, H., Wallace, B. C. and Bau, D. (2026) 'Using OCR Heads to Verbalize Image Semantics', Available at: https://omanscience.com/en/articles/using-ocr-heads-to-verbalize-image-semantics.

Vancouver

Feucht S, Krojer B, Wang S, Abrahamsen H, Wallace BC, Bau D. Using OCR Heads to Verbalize Image Semantics. https://omanscience.com/en/articles/using-ocr-heads-to-verbalize-image-semantics

IEEE

S. Feucht, B. Krojer, S. Wang, H. Abrahamsen, B. C. Wallace, and D. Bau, "Using OCR Heads to Verbalize Image Semantics," https://omanscience.com/en/articles/using-ocr-heads-to-verbalize-image-semantics.