Preprint Open access
InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning
Table images preserve structural information that are often lost in text serialization, and reasoning over them requires locating relevant rows, columns, and cells step by step. Current multimodal large language models (MLLMs) encode the whole image once before reasoning, so they cannot pick up row-, column-, and cell- …