Preprint Open access
Where Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models
Unified vision-language models (VLMs) that tokenize images through a vector-quantized (VQ) codebook routinely hallucinate objects on grounded yes/no benchmarks, yet existing decoding-time fixes treat this as generic miscalibration without an architectural account. Using activation patching across twenty-five models spa …