الباحثون

Yunfei Chu

المنشورات 4

نسخة أولية وصول مفتوح

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

Yuan Feng, Qize Yang, Ruizhe Chen وآخرون · 2026

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Down …

نسخة أولية وصول مفتوح

MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training

Yuxin Liu, Yuxuan Wang, Zhenxin Lei وآخرون · 2026

Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear. We introduce MMPostTrainBench, a benchmark spanning eight t …

نسخة أولية وصول مفتوح

OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

Junming Lin, Yuxuan Wang, Zhenxin Lei وآخرون · 2026

Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We addre …

نسخة أولية وصول مفتوح

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

Qi Chen, Yunfei Chu, Haolin He وآخرون · 2026

Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental qu …

المؤلفون المشاركون