الباحثون

Yu-Gang Jiang

المنشورات 8

نسخة أولية وصول مفتوح

OpenViTac: Learning and Benchmarking Visuo-Tactile Policies in a Unified Sim-and-Real Framework

Yifan Wu, Qin Li, Nan Min وآخرون · 2026

Tactile feedback provides embodied agents with physical information beyond visual observations, enabling more reliable interaction with the real world. However, despite the rapid progress of vision-tactile-language-action (VTLA) policies, there remains a lack of unified benchmarks for evaluating tactile-enabled robot m …

نسخة أولية وصول مفتوح

Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training

Shuyuan Tu, Qi Tian, Yinming Huang وآخرون · 2026

Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, diluting learning signals for informativ …

نسخة أولية وصول مفتوح

LIBERO-Agent: Evaluating General-Purpose Agents for Direct Embodied Manipulation

Zijie Diao, Yitong Chen, Sicheng Xie وآخرون · 2026

General-purpose agents can plan, use tools, and revise their behavior from feedback, but it remains unclear whether these capabilities transfer from digital environments to embodied manipulation. To investigate this question, we introduce LIBERO-Agent, an agent-native benchmark for evaluating these agents in robot mani …

نسخة أولية وصول مفتوح

Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents

Sicheng Xie, Yitong Chen, Haidong Cao وآخرون · 2026

Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by …

نسخة أولية وصول مفتوح

MaLiang-Harness: A Programmable Path to Image and Video Generation

Haoyu Zhao, Zihao Zhang, Xudong Wang وآخرون · 2026

Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap an …

نسخة أولية وصول مفتوح

LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models

Huiqiong Li, Zhiting Mei, Anirudha Majumdar وآخرون · 2026

Robotic foundation models achieve impressive performance on standard manipulation benchmarks, yet these evaluations typically assume clean, timely, and consistent visual observations throughout execution. We introduce LIBERO-VPro, a benchmark for systematically evaluating the closed-loop visual robustness of robotic fo …

نسخة أولية وصول مفتوح

Stable and Efficient Real-World Online VLA Post-Training via Asynchronous Replay-Anchored Policy Improvement

Jiarui Yang, Jiajin Zhang, Bin Zhu وآخرون · 2026

Online post-training of vision-language-action (VLA) models requires efficient use of robot interaction and reliable policy improvement from continually collected experience. We propose asynchronous Replay-Anchored Policy improvement (RAPolicy), a framework that performs rollout and learning concurrently while groundin …

نسخة أولية وصول مفتوح

Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands

Zhenjie Yang, Yideng Zhang, Dongjie Zhang وآخرون · 2026

Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not converged to a common design. Dexterous hands differ in finger structure, contact surfaces, and sensor layouts, while simulated tactile signals still differ from measurements p …

المؤلفون المشاركون