Abstract

Studying the alignment between the internal representations of vision models and the responses of the visual cortex to the same observed visual stimuli has enabled us to better understand human visual processing. However, studies so far have largely overlooked the fact that the human brain not only processes observed visual stimuli, but also predicts upcoming stimuli based on what has been observed. Accordingly, we hypothesize that internal representations for generating future video frames are better aligned with the predictive nature of human visual processing than representations of the observed video itself. To this end, we compare the alignment between human video-watching fMRI responses in the visual cortex and the internal representations from two types of video diffusion models, an autoregressive (AR) model and its non-AR base model. We first conduct a within-model analysis of the AR video diffusion model and show that the representations for future video generation align better with the visual cortex than the representations of the observed video. We then compare the internal representations of the AR model with those of its non-AR base model and again show that the representations for future video generation align better with the visual cortex than the representations for observed video reconstruction by the base model. Specifically, the alignment of observed video reconstruction is concentrated in lower-order visual cortex, whereas that of future video generation is concentrated in higher-order visual cortex. Finally, we show in a human behavioral experiment that humans prefer videos generated by amplifying the contributions of individual layers that align better with the visual cortex.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Bang, C. B., Chung, H., & Kim, B. H. (2026). Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video. https://omanscience.com/en/articles/future-video-generation-better-aligns-with-the-human-visual-cortex-than-observed-video

MLA 9

Bang, Chang-Bae, et al. "Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video." https://omanscience.com/en/articles/future-video-generation-better-aligns-with-the-human-visual-cortex-than-observed-video.

Chicago (author–date)

Bang, Chang-Bae, Hyungjin Chung, and Byung-Hoon Kim. 2026. "Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video." https://omanscience.com/en/articles/future-video-generation-better-aligns-with-the-human-visual-cortex-than-observed-video.

Harvard

Bang, C. B., Chung, H. and Kim, B. H. (2026) 'Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video', Available at: https://omanscience.com/en/articles/future-video-generation-better-aligns-with-the-human-visual-cortex-than-observed-video.

Vancouver

Bang CB, Chung H, Kim BH. Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video. https://omanscience.com/en/articles/future-video-generation-better-aligns-with-the-human-visual-cortex-than-observed-video

IEEE

C. B. Bang, H. Chung, and B. H. Kim, "Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video," https://omanscience.com/en/articles/future-video-generation-better-aligns-with-the-human-visual-cortex-than-observed-video.