الباحثون

Yusuke Iwasawa

المنشورات 5

نسخة أولية وصول مفتوح

Rendering-Free Lookahead for Question-Guided Active Vision

Koya Sakamoto, Daichi Azuma, Shuhei Kurita وآخرون · 2026

Active robot vision requires controlling the camera to reveal task-relevant information that is hidden from the current viewpoint. For example, determining what is inside a box may require raising the camera and looking down into it. For viewpoint-dependent question answering, the challenge is to select camera motions …

نسخة أولية وصول مفتوح

Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs

Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning. Existing benchmarks document this failure but cannot say \emph{why} it happens or which cognitive capability is missing. We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm fr …

نسخة أولية وصول مفتوح

PHASE: Compliance-Enabled Tactile Phase Retrieval for Few-Shot Insertion Learning

Contact-rich assembly tasks such as peg-in-hole insertion remain difficult to learn from limited demonstrations. While retrieval-augmented imitation learning, which augments target demonstrations with relevant prior data, offers a promising direction, its applicability to contact-rich manipulation remains largely unexp …

نسخة أولية وصول مفتوح

Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision

As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action models (LAMs) address both by learning latent actions from action-free videos that can be shared across embodiments, however, in practice, LAM …

المؤلفون المشاركون