الباحثون

Steve Yves

المنشورات 2

نسخة أولية وصول مفتوح

OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning

Zhongyu Yang, Jiale Tao, Ruitao Chen وآخرون · 2026

Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide covera …

نسخة أولية وصول مفتوح

AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

Yuxiang Wang, Kunyu Feng, Yuancheng Wang وآخرون · 2026

Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic …

المؤلفون المشاركون