الباحثون

Shaoyu Chen

المنشورات 2

نسخة أولية وصول مفتوح

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces

Hongyuan Tao, Xinggang Wang, Lianghui Zhu وآخرون · 2026

We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with continuous image generation. The former introduces a visual quantization bottleneck. The …

نسخة أولية وصول مفتوح

ReDrive: Shaping Representations with World Modeling for End-to-End Driving

Yueting Zhu, Shaoyu Chen, Yuehao Song وآخرون · 2026

Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system architecture. Inspired by …

المؤلفون المشاركون