الباحثون

Xinggang Wang

المنشورات 6

نسخة أولية وصول مفتوح

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces

Hongyuan Tao, Xinggang Wang, Lianghui Zhu وآخرون · 2026

We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with continuous image generation. The former introduces a visual quantization bottleneck. The …

نسخة أولية وصول مفتوح

Rethinking Representations for World-Action Modeling

Haoyi Jiang, Liu Liu, Xinjiang Wang وآخرون · 2026

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure e …

نسخة أولية وصول مفتوح

ReDrive: Shaping Representations with World Modeling for End-to-End Driving

Yueting Zhu, Shaoyu Chen, Yuehao Song وآخرون · 2026

Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system architecture. Inspired by …

نسخة أولية وصول مفتوح

ForeDrive: Foresight-Guided End-to-End Autonomous Driving with a Planning-Relevant Latent World Model

Sinuo Wang, Zichong Gu, Yuhan Huang وآخرون · 2026

Existing latent world models are typically optimized for future predictability, yet the resulting representations are not necessarily useful for planning in autonomous driving. Predictions are commonly used for pretraining or auxiliary supervision rather than as direct conditioning signals for trajectory generation. We …

المؤلفون المشاركون