الباحثون

Haolin He

المنشورات 2

نسخة أولية وصول مفتوح

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

Yuan Feng, Qize Yang, Ruizhe Chen وآخرون · 2026

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Down …

نسخة أولية وصول مفتوح

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

Qi Chen, Yunfei Chu, Haolin He وآخرون · 2026

Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental qu …

المؤلفون المشاركون