الباحثون

Kai Yu

المنشورات 2

نسخة أولية وصول مفتوح

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces

Hongyuan Tao, Xinggang Wang, Lianghui Zhu وآخرون · 2026

We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with continuous image generation. The former introduces a visual quantization bottleneck. The …

نسخة أولية وصول مفتوح

MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

Mei Wu, Rui Xie, Runyu Zhang وآخرون · 2026

Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-p …

المؤلفون المشاركون