Abstract

Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them. We study this problem through multi-image re-representation, viewing prompted Chain-of-Thought reasoning and agentic visual tool use as different ways of re-organising visual evidence during reasoning. We introduce Mosaic, a general-purpose multi-image visual harness that enables an MLLM to actively construct visual intermediates with ten composable image operations. We compare five re-representation settings on existing multi-image benchmarks and on MosaicBench, a new grounding-focused benchmark for fine-grained multi-image understanding. Our experiments show that the relative benefits of textual and visual re-representation are strongly task-dependent. Visual re-representation is particularly effective for tasks requiring precise visual evidence, including hypothesis testing, precision comparison, and orientation-sensitive reasoning, while tasks dominated by higher-level semantic content show smaller or less consistent gains. Building on this finding, we train MosaicAgent-8B to use Mosaic with reinforcement learning using only accuracy and format rewards. Without demonstration trajectories or rewards for specific tool-use, the agent learns to compose visual operations over multiple steps and exhibits diverse problem-solving patterns unpromptedly. Code and data will be released at https://github.com/gengyuanmax/Mosaic.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, G., Han, X., Xie, X., Liu, T., & Tresp, V. (2026). Rethinking Multi-Image Re-Representation in Multi-Image Understanding. https://omanscience.com/en/articles/rethinking-multi-image-re-representation-in-multi-image-understanding

MLA 9

Zhang, Gengyuan, et al. "Rethinking Multi-Image Re-Representation in Multi-Image Understanding." https://omanscience.com/en/articles/rethinking-multi-image-re-representation-in-multi-image-understanding.

Chicago (author–date)

Zhang, Gengyuan, Xiao Han, Xinyu Xie, Tong Liu, and Volker Tresp. 2026. "Rethinking Multi-Image Re-Representation in Multi-Image Understanding." https://omanscience.com/en/articles/rethinking-multi-image-re-representation-in-multi-image-understanding.

Harvard

Zhang, G., Han, X., Xie, X., Liu, T. and Tresp, V. (2026) 'Rethinking Multi-Image Re-Representation in Multi-Image Understanding', Available at: https://omanscience.com/en/articles/rethinking-multi-image-re-representation-in-multi-image-understanding.

Vancouver

Zhang G, Han X, Xie X, Liu T, Tresp V. Rethinking Multi-Image Re-Representation in Multi-Image Understanding. https://omanscience.com/en/articles/rethinking-multi-image-re-representation-in-multi-image-understanding

IEEE

G. Zhang, X. Han, X. Xie, T. Liu, and V. Tresp, "Rethinking Multi-Image Re-Representation in Multi-Image Understanding," https://omanscience.com/en/articles/rethinking-multi-image-re-representation-in-multi-image-understanding.