الباحثون

Di Wu

المنشورات 10

نسخة أولية وصول مفتوح

Execution-Aligned Progressive Noise for Consistent Asynchronous Replanning in Generative Robot Policies

Di Wu, Ping Liu, Xuhua Chen وآخرون · 2026

Continuous asynchronous replanning is essential for real-time generative robot policies, but independent stochastic initialization can cause mode switching and inconsistent continuation across action chunks. We propose Execution-Aligned Progressive Noise (EAPN), which introduces structured stochasticity at both inter-c …

نسخة أولية وصول مفتوح

ColdDDI: Evaluating Knowledge Utilization in Cold-Start Drug-Drug Interaction Prediction

Jiheng Liang, Chen Zhao, Di Wu وآخرون · 2026

Cold-start drug-drug interaction (DDI) prediction tests whether models can identify clinically significant interactions for drugs without training-time interaction history. Existing benchmarks mostly report aggregate edge-prediction scores, leaving a key evaluation question unanswered: when models receive molecular, te …

نسخة أولية وصول مفتوح

Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation

Di Wu, Rongtian Shen, Ping Liu وآخرون · 2026

Vision-language-action (VLA) models face a timing gap between low-rate inference and high-rate robot execution. We characterize this gap through end-to-end latency measurements of model inference and the robot execution chain. Repeated Flow Matching denoising contributes substantially to inference cost, while robot-sid …

نسخة أولية وصول مفتوح

Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence

Xuhua Chen, Zhenhan Yin, Yuan Zhang وآخرون · 2026

World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0 …

نسخة أولية وصول مفتوح

Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools

Shijia Ge, Alex Zhou, Jianshu Zeng وآخرون · 2026

Frontier models can control robots, but reasoning through every reach, grasp, and retreat makes manipulation slow and token-intensive. We revisit code as policy with a different division of labor: models build executable tools, code handles multi-phase motions, and models decide what to do next. We introduce URAI (Univ …

نسخة أولية وصول مفتوح

AgBench: Agentic AI Benchmarks for Personal AI Devices

Yizhou Han, Di Wu, Dhananjay Saikumar وآخرون · 2026

Agentic AI systems increasingly rely on cloud-hosted large language models for planning, tool use, and iterative execution, raising concerns about API cost and data exposure. Advances in personal AI devices enable agents to execute locally, but limited resources on device may affect task success and performance. Existi …

نسخة أولية وصول مفتوح

Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence

Hongcheng Gao, Jingjing Zhou, Zelin Zheng وآخرون · 2026

Vision-language-action (VLA) and world-action (WAM) models map observations and instructions directly to robot actions. This directness ties a policy to training: minor layout or viewpoint changes cause failure, and instructions generalize poorly. The root cause lies in representation: task requirements, conditions, pr …

نسخة أولية وصول مفتوح

AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining

Di Wu, Dongchen Zheng, Junhe Sheng وآخرون · 2026

Embodied foundation models are constrained by the limited scale and diversity of robot demonstrations, motivating the use of large-scale egocentric human interaction data. However, how to effectively incorporate such data into embodied-model pre-training remains unclear because of substantial embodiment and action-spac …

نسخة أولية وصول مفتوح

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI, Anyi Xu, B. Li وآخرون · 2026

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Togeth …

المؤلفون المشاركون