Authors

Xiangbo Gao

Publications 3

Preprint Open access

Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy

Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics …

Preprint Open access

AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation

Keshu Wu, Hao Zhang, Rui Gan et al. · 2026

Air-ground transportation research increasingly relies on co-simulation, yet constructing scenarios remains labor-intensive and difficult to validate. More importantly, a generated scenario may execute successfully while failing to realize the spatial, temporal, communication, or behavioral relationships requested by t …

Co-authors