نسخة أولية وصول مفتوح
Medical Image Alignment Assessment as a Test of Generalist Visual Reasoning in Frontier Multimodal Models
Frontier multimodal large language models (MLLMs) are increasingly positioned as general purpose visual reasoners as part of the quest for artificial general intelligence. A key test of this generality is whether they can perform novel visual judgments that humans can make reliably from visual evidence and task instruc …