الباحثون

Yi Yang

المنشورات 6

نسخة أولية وصول مفتوح

WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models

Kai Ding, Yang He, Ruijie Quan وآخرون · 2026

World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT). In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert …

نسخة أولية وصول مفتوح

Point It, Strike It: Direction-Conditioned Dynamic Manipulation of Deformable Linear Objects

Yi Yang, Xiang Fei, Lehong Wang وآخرون · 2026

Goal-conditioned dynamic manipulation of deformable linear objects has mainly specified goals as positions for a rope tip to reach. Many tasks, however, depend on how the tip arrives. We therefore study single-swing rope striking with goals that specify the tip's 3D position and arrival direction, across the workspace …

نسخة أولية وصول مفتوح

On-Policy Distillation Teaches New Skills but Not New Knowledge

Yixuan Tang, Yi Yang · 2026

On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently …

نسخة أولية وصول مفتوح

GRC-Pose: Generation-Reconstruction Correspondence for Prior-Free 6D Object Pose Tracking

Shiyang Liu, Weiquan Lin, Luping Xiao وآخرون · 2026

Prior-free 6D object pose tracking seeks to recover the trajectory of an unseen object from a single RGB video without object-specific CAD models, posed reference images, or pose annotations. Geometric foundation models provide complementary object-centric and scene-centric cues, yet SAM3D CAD is indexed by an arbitrar …

نسخة أولية وصول مفتوح

RAVEL: Asynchronous Rolling Inference for Flow-Based Vision-Language-Action Models

Yuhan Chen, Ke Yu, Pengfei Liu وآخرون · 2026

Flow-based vision-language-action (VLA) models are highly effective for generalist robot manipulation, yet their reliance on computationally expensive VLM encoding and multi-step iterative action generation imposes a significant latency bottleneck. The resulting inference latency makes it difficult for robots to respon …

نسخة أولية وصول مفتوح

TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models

Shuning Wang, Zhiheng Wu, Xun Zhou وآخرون · 2026

Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reas …

المؤلفون المشاركون