Preprint Open access
OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide covera …