Preprint Open access
From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual struc …