Abstract

Audio-video generation using heterogeneous multimodal references has emerged as a new challenge, requiring both compositional control over generation and grounded understanding of multimodal context. In this paper, we introduce ORAV Bench for Omni Reference Audio-Video Generation, comprising 380 task instances with 2-10 references, 9 semantic roles, and 30 role compositions. Instructions specify the relationships among references; the media supply the identities, dynamics, and audio characteristics to be realized. To evaluate these open-ended outputs, we develop a reference-aware pairwise protocol that prepares visual and auditory evidence, compares the intended contribution of each reference, and checks the overall verdict in both presentation orders. On held-out instances, it achieves 86.08% effective agreement with human judgments. Across 5 frontier systems, overall rankings conceal distinct strengths across reference compositions. A recurring failure is to reproduce unintended source content in place of the requested result, despite closely resembling a reference. Reproducible pointwise diagnostics of quality, reference affinity, and speech reveal distinct dimensions of model behavior. ORAV thus offers a benchmark for tracking progress toward controllable, compositional, and reference-faithful audio-video generation.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Hua, J., Feng, X., Hua, J., Liu, C., Wang, B., & Liu, M. (2026). ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts. https://omanscience.com/en/articles/orav-benchmarking-audio-video-generation-from-multimodal-contexts

MLA 9

Hua, Jiacheng, et al. "ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts." https://omanscience.com/en/articles/orav-benchmarking-audio-video-generation-from-multimodal-contexts.

Chicago (author–date)

Hua, Jiacheng, Xiaokun Feng, Jiaqi Hua, Chang Liu, Biao Wang, and Miao Liu. 2026. "ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts." https://omanscience.com/en/articles/orav-benchmarking-audio-video-generation-from-multimodal-contexts.

Harvard

Hua, J., Feng, X., Hua, J., Liu, C., Wang, B. and Liu, M. (2026) 'ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts', Available at: https://omanscience.com/en/articles/orav-benchmarking-audio-video-generation-from-multimodal-contexts.

Vancouver

Hua J, Feng X, Hua J, Liu C, Wang B, Liu M. ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts. https://omanscience.com/en/articles/orav-benchmarking-audio-video-generation-from-multimodal-contexts

IEEE

J. Hua, X. Feng, J. Hua, C. Liu, B. Wang, and M. Liu, "ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts," https://omanscience.com/en/articles/orav-benchmarking-audio-video-generation-from-multimodal-contexts.