نسخة أولية وصول مفتوح
Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs
Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning. Existing benchmarks document this failure but cannot say \emph{why} it happens or which cognitive capability is missing. We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm fr …