الباحثون

Xize Cheng

المنشورات 3

نسخة أولية وصول مفتوح

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

Yuan Feng, Qize Yang, Ruizhe Chen وآخرون · 2026

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Down …

نسخة أولية وصول مفتوح

MuLA-Bench: A Multilingual Long-Form Audio Understanding Benchmark via Multi-Tier Auditing

Zeyu Yang, Xinyu Zhang, Zibo Bi وآخرون · 2026

Long-form audio performance is often summarized by context length and aggregate accuracy, obscuring how language, evidence, and task jointly shape difficulty. We introduce MuLA-Bench: 5,038 open-ended questions over 1,769 in-the-wild recordings totaling 1,377.9 hours, covering 16 languages and eight domains. A balanced …

نسخة أولية وصول مفتوح

Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models

Xize Cheng, Wenxu Jia, Chenyuhao Wen وآخرون · 2026

Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient information. This issue is even more pronounced in the speech domain, where audio, as a low-compression modality, requires substantially more embeddings than text to prese …

المؤلفون المشاركون