الباحثون

Qifeng Chen

المنشورات 4

نسخة أولية وصول مفتوح

UniData: Universal Multimodal Instruction Generation Pipeline

Jiaqi Tang, Yi-Feng Wu, Yuting Zhang وآخرون · 2026

Multimodal Large Language Models (MLLMs) are increasingly being applied in a wider range of real-world scenarios. However, due to the substantial labor cost, creating high-quality multimodal instruction datasets for MLLMs remains a significant challenge. Although some methods propose to generate instruction data, they …

نسخة أولية وصول مفتوح

WorldSonus: Bringing Sound to Worlds

Pengjun Fang, Jingyi Fa, Kam Man Wu وآخرون · 2026

Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream so …

نسخة أولية وصول مفتوح

HelixWorld: A Real-time Interactive Audio-Visual World Model

Lei Ke, Jiahao Pan, Zeyue Tian وآخرون · 2026

World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive aud …

نسخة أولية وصول مفتوح

The Past Frames the Future: Memory for Autoregressive Video Generation

Harold Haodong Chen, Rongjin Guo, Disen Lan وآخرون · 2026

Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands, …

المؤلفون المشاركون