الباحثون

Zuxuan Wu

المنشورات 8

نسخة أولية وصول مفتوح

OpenViTac: Learning and Benchmarking Visuo-Tactile Policies in a Unified Sim-and-Real Framework

Yifan Wu, Qin Li, Nan Min وآخرون · 2026

Tactile feedback provides embodied agents with physical information beyond visual observations, enabling more reliable interaction with the real world. However, despite the rapid progress of vision-tactile-language-action (VTLA) policies, there remains a lack of unified benchmarks for evaluating tactile-enabled robot m …

نسخة أولية وصول مفتوح

Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training

Shuyuan Tu, Qi Tian, Yinming Huang وآخرون · 2026

Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, diluting learning signals for informativ …

نسخة أولية وصول مفتوح

LIBERO-Agent: Evaluating General-Purpose Agents for Direct Embodied Manipulation

Zijie Diao, Yitong Chen, Sicheng Xie وآخرون · 2026

General-purpose agents can plan, use tools, and revise their behavior from feedback, but it remains unclear whether these capabilities transfer from digital environments to embodied manipulation. To investigate this question, we introduce LIBERO-Agent, an agent-native benchmark for evaluating these agents in robot mani …

نسخة أولية وصول مفتوح

Chinese-Jev: Bringing System One Model to Chinese-Language Tasks

Zexiao Wang, Zihao Zhang, Xudong Wang وآخرون · 2026

System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper …

نسخة أولية وصول مفتوح

Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents

Sicheng Xie, Yitong Chen, Haidong Cao وآخرون · 2026

Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by …

نسخة أولية وصول مفتوح

MaLiang-Harness: A Programmable Path to Image and Video Generation

Haoyu Zhao, Zihao Zhang, Xudong Wang وآخرون · 2026

Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap an …

نسخة أولية وصول مفتوح

Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands

Zhenjie Yang, Yideng Zhang, Dongjie Zhang وآخرون · 2026

Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not converged to a common design. Dexterous hands differ in finger structure, contact surfaces, and sensor layouts, while simulated tactile signals still differ from measurements p …

المؤلفون المشاركون