الباحثون

Kai Wang

المنشورات 14

نسخة أولية وصول مفتوح

OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning

Zhongyu Yang, Jiale Tao, Ruitao Chen وآخرون · 2026

Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide covera …

نسخة أولية وصول مفتوح

FaceKit: a Toolkit for Interpretable Facial Phenotyping, Synthetic Image Generation and Privacy Analysis in Rare Diseases

Many rare genetic diseases are associated with recognizable craniofacial features. However, traditional approaches for describing facial morphology rely largely on qualitative clinical observation and free-text descriptions, which are often subjective, non-standardized, and difficult to reproduce across observers and i …

نسخة أولية وصول مفتوح

ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models

Yuan Xu, Yixiang Chen, Qisen Ma وآخرون · 2026

Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the vis …

نسخة أولية وصول مفتوح

Flash-OPD: Fast On-Policy Distillation

Wei Chen, Junle Chen, Yitong Yang وآخرون · 2026

On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but generating and evaluating long rollouts incurs substantial training cost. Existing acceleration methods reduce this cost through open-loop rollout schedules or closed-loop horizon adaptation. However, supervision comp …

نسخة أولية وصول مفتوح

You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn Dialogue

Junle Chen, Wei Chen, Zhengjun Huang وآخرون · 2026

When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed. We find a surprising failure: merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged. To systematical …

نسخة أولية وصول مفتوح

DiFF: Doppler-informed Flow Matching for Human Motion Flow

Perceiving human motion via privacy-preserving 4D millimeter-wave (mmWave) radar is critical for next-generation human-robot interaction (HRI), where point cloud scene flow serves as a foundational motion representation. Yet the extreme sparsity and noise of 4D radar point clouds make non-rigid motion flow estimation s …

نسخة أولية وصول مفتوح

BTC3D: Blended Tile Conditioning for Detail-Enhancing Image-to-3D Generation

Junyu Li, Qiuyu Chen, Pengcheng Wang وآخرون · 2026

Recent diffusion-based pipelines have achieved promising progress in image-to-3D synthesis. However, generating high-fidelity details remains challenging, especially when the input image contains rich details. Existing approaches often rely on globally encoded conditioning features, which compress spatial information a …

المؤلفون المشاركون