الباحثون

Xiaodan Liang

المنشورات 3

نسخة أولية وصول مفتوح

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu وآخرون · 2026

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert. We present EWAM, an action-centric unified embo …

نسخة أولية وصول مفتوح

VR-JEPA: Learning Contrastive-State Latent Guidance for Generation-based Video Reasoning

Zehua Ma, Kun Xiang, Yunshuang Nie وآخرون · 2026

Reasoning through video generation offers a promising path toward visual intelligence by modeling latent visual states and their dynamics. However, current video generation models often lack explicit guidance on how these states should evolve, leaving generated trajectories prone to physical and structural inconsistenc …

نسخة أولية وصول مفتوح

Mind the RefGAP: Correcting Reference Attention in Diffusion-Based Visual Editing

Yanan Wang, Shengcai Liao, Guangyi Liu وآخرون · 2026

Reference-guided diffusion editors struggle to faithfully reproduce user-provided references. We identify a potential bottleneck in diffusion editors: many methods provide limited reference-attention allocation. For example, in LoomVideo, edit-region queries assign less than 1% of their attention mass to the reference. …

المؤلفون المشاركون