الباحثون

Ling Wang

المنشورات 4

نسخة أولية وصول مفتوح

OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

Junming Lin, Yuxuan Wang, Zhenxin Lei وآخرون · 2026

Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We addre …

نسخة أولية وصول مفتوح

From Pixel Generation to Topological Inference: Structural Dual Super-Resolution for Trustworthy Cross-Physical-Domain Trabecular Morphology Learning

Clinical CT and UHRCT cannot resolve individual trabeculae, whereas synchrotron radiation microCT (SRuCT) provides high-resolution references but is not applicable for in vivo imaging. The two domains differ by a 32x resolution gap, are only coarsely paired, and exhibit severe physical differences including partial vol …

نسخة أولية وصول مفتوح

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

Qi Chen, Yunfei Chu, Haolin He وآخرون · 2026

Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental qu …

نسخة أولية وصول مفتوح

NavPatch: Evidence-Guided Object-Level Costmap Correction with Vision-Language Models

Shiji Sun, Xingyu Tao, Hao Wang وآخرون · 2026

Mobile robots typically rely on geometric maps for obstacle avoidance and path planning, but the resulting obstacle representation does not always match how an object should affect navigation. A low lying cable may be missed, a flexible curtain may create spurious blockage, and a traffic cone may require an exclusion r …

المؤلفون المشاركون