الباحثون

Hiroki Furuta

المنشورات 3

نسخة أولية وصول مفتوح

Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs

Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning. Existing benchmarks document this failure but cannot say \emph{why} it happens or which cognitive capability is missing. We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm fr …

نسخة أولية وصول مفتوح

VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation

Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple refer …

المؤلفون المشاركون