الباحثون

Tatsuya Matsushima

المنشورات 4

نسخة أولية وصول مفتوح

RoboPace: Contact-Aware Time-Optimal Retiming for Action-Chunk Policies

Robot manipulation data collection has been shifting from teleoperation toward robot-free demonstrations, through interfaces such as the Universal Manipulation Interface (UMI) or directly from human hands. Vision-Language-Action (VLA) policies trained on such data inherit the demonstrator's timing. Yet human timing doe …

نسخة أولية وصول مفتوح

YUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language Grounding

Vision-Language-Action (VLA) models acquire broad manipulation capabilities via large-scale pretraining, yet eliciting them through language requires fine-grained alignment between instructions and physical interactions. Existing robot demonstrations typically provide only coarse task descriptions, omitting how actions …

نسخة أولية وصول مفتوح

PHASE: Compliance-Enabled Tactile Phase Retrieval for Few-Shot Insertion Learning

Contact-rich assembly tasks such as peg-in-hole insertion remain difficult to learn from limited demonstrations. While retrieval-augmented imitation learning, which augments target demonstrations with relevant prior data, offers a promising direction, its applicability to contact-rich manipulation remains largely unexp …

نسخة أولية وصول مفتوح

Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision

As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action models (LAMs) address both by learning latent actions from action-free videos that can be shared across embodiments, however, in practice, LAM …

المؤلفون المشاركون