Authors

Yilun Du

Publications 7

Preprint Open access

Robot Learning with Visual Predicted Force

Haonan Chen, Feiyang Wu, Yuxiang Ma et al. · 2026

Force-aware manipulation typically relies on specialized force or tactile sensors. We show that force-aware manipulation can instead be achieved through visual force prediction from the deformation of a compliant Fin Ray gripper. Our approach trains two models. First, we train a visual force estimator on calibration da …

Preprint Open access

MoSE3: Learning World-Space SE(3) at Every Pixel

Jiahuan Cheng, Zhiyi Li, Tian Xia et al. · 2026

Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model tha …

Preprint Open access

Finetuning with Sampling: SFT Learns Better Than You Think

Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is …

Preprint Open access

Training Object Permanence in World Models

Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged …

Preprint Open access

Representation-guided in-context learning for medical image interpretation with multimodal large language models

Minda Zhao, Fangyu Hu, Yan Luo et al. · 2026

Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires resource-intensive domain-specific fine-tuning. Here we introduce representation-guided in-context learning (RG-ICL), a training-free inference framework that retrieves que …

Co-authors