الباحثون

Yiming Li

المنشورات 9

نسخة أولية وصول مفتوح

Salvation Lies Within: Eliciting Inherent Style Transfer in Step-Distilled Diffusion Models

Shengyin Sun, Yiming Li, Yingzhao Lian وآخرون · 2026

Adapting step-distilled text-to-image (T2I) models through post-training incurs additional computational costs and affects native few-step generation behavior. This motivates a complementary route beyond style-specific adaptation: drawing on the visual knowledge already encoded in step-distilled T2I models to elicit st …

نسخة أولية وصول مفتوح

Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs

Xiaojun Jia, Simeng Qin, Yiming Li وآخرون · 2026

Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples, especially in black-box settings where only open-source surrogate models are accessible. Existing targeted transfer attacks mainly align adversarial and target samples using global image-level features, such as encoder [CLS …

نسخة أولية وصول مفتوح

Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation

Sichao Liu, Zekun Wang, Lixuan Tang وآخرون · 2026

Robotic manipulation policies are advancing rapidly with increasing reliance on vision-language models for end-to-end decision making. However, reliable deployment remains challenging because many policies lack explicit mechanisms for predicting task outcomes and evaluating whether generated actions will achieve desire …

نسخة أولية وصول مفتوح

TactileStep: Sole Tactile Learning for Regulating Foot-Terrain Interaction in Humanoid Locomotion

Zizhuo Wang, Ming-Ju Lee, Shaoting Zhu وآخرون · 2026

Humanoid parkour policies can traverse various terrains, but task completion may mask challenges of harsh landings, edge contacts, and unstable stance contacts. Humans naturally regulate foot-terrain interaction through tactile feedback, modulating contact compliance according to terrain stiffness. This highlights a ke …

نسخة أولية وصول مفتوح

DAVIS: A Depth-Only End-to-End Active-Vision Framework for Humanoid Soccer Skills

Jiakang Jin, Yixiao Huo, Pengyuan Wang وآخرون · 2026

Humanoid soccer contact skills require more than producing high-impact foot-ball contacts: the robot must close the loop over perception, approach, alignment, impact, and recovery while its own motion induces substantial viewpoint changes, frequent loss of the ball from view, and uncertain contact outcomes. In this wor …

نسخة أولية وصول مفتوح

SAMI3D-DW: Interactive Segmentation of Any 3D Medical Images

Ping Gong, Shiyuan Su, Fandong Zhang وآخرون · 2026

Interactive segmentation of 3D medical images supports quantitative analysis of anatomical structures and disease while allowing users to specify and refine their targets. Despite substantial progress by nnInteractive and VISTA3D, reliable segmentation across diverse clinical targets remains challenging, particularly f …

نسخة أولية وصول مفتوح

Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation

Lijie Wang, Zheng Lu, Yiming Wang وآخرون · 2026

Vision-Language-Action models and World-Action Models have advanced language-conditioned robotic manipulation, yet often leave metric relations among actions, objects, and scene geometry implicit. Human manipulation combines semantic understanding of task-relevant objects with spatial feedback that guides hand motion r …

المؤلفون المشاركون