الباحثون

Yutaka Matsuo

المنشورات 6

نسخة أولية وصول مفتوح

Rendering-Free Lookahead for Question-Guided Active Vision

Koya Sakamoto, Daichi Azuma, Shuhei Kurita وآخرون · 2026

Active robot vision requires controlling the camera to reveal task-relevant information that is hidden from the current viewpoint. For example, determining what is inside a box may require raising the camera and looking down into it. For viewpoint-dependent question answering, the challenge is to select camera motions …

نسخة أولية وصول مفتوح

Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs

Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning. Existing benchmarks document this failure but cannot say \emph{why} it happens or which cognitive capability is missing. We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm fr …

نسخة أولية وصول مفتوح

SLIM: Simplex-Lattice Interpolation Merging

Optimizing merging coefficients for large language models can require many costly benchmark evaluations. We propose \textbf{Simplex-Lattice Interpolation Merging (SLIM)}, which constructs a quadratic surrogate of aggregate performance on the coefficient simplex using a classical mixture design. Evaluations of individua …

نسخة أولية وصول مفتوح

VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation

Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple refer …

نسخة أولية وصول مفتوح

Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision

As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action models (LAMs) address both by learning latent actions from action-free videos that can be shared across embodiments, however, in practice, LAM …

المؤلفون المشاركون