الباحثون

Yuzhang Shang

المنشورات 7

نسخة أولية وصول مفتوح

Keepsake: Selective Spatial Memory for Long-Horizon Video Generation

Long-horizon camera-controlled video generation relies on persistent memory to maintain scene consistency. Existing systems follow two strategies to achieve this consistency. Full-history approaches retain all generated observations, causing unbounded storage and retrieval costs. Selective-construction approaches reduc …

نسخة أولية وصول مفتوح

Recursive Video In-Context Learning for Agentic Robot

Wenrui Bao, Xinxin Liu, Bingxin Xu وآخرون · 2026

LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact …

نسخة أولية وصول مفتوح

CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment

Kunyang Li, Hai Nguyen, Joshua Lowe وآخرون · 2026

Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial ma …

نسخة أولية وصول مفتوح

PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration

Dannong Wang, Yuran Zhang, Bian Sun وآخرون · 2026

Clinical large language model (LLM) agents deployed locally can consult more capable remote models, but doing so risks exposing patient information. Privacy-conscious delegation places disclosure decisions with a local agent, yet removing explicit identifiers is insufficient: quasi-identifiers can accumulate across mul …

نسخة أولية وصول مفتوح

DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model

Tianyun Jiang, Wenrui Bao, Bingxin Xu وآخرون · 2026

World-Action Models (WAMs) offer visual foresight for robotic manipulation, but pixel-space models repeatedly reconstruct entire future scenes, incurring high computational cost and spatio-temporal redundancy. In physical manipulation, consecutive frames often share most of their visual context; the changes between the …

نسخة أولية وصول مفتوح

Coding Agents with Harness for Safe Robot Control

Bingxin Xu, Yuzhang Shang, Zhen Dong وآخرون · 2026

Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the robot controller as a program, and agents built in this way now operate robots without robot-specific training. Whether this paradigm is also safe, however, has not been asked. We evaluate coding agents under a safety …

المؤلفون المشاركون