الباحثون

Cheng Deng

المنشورات 2

نسخة أولية وصول مفتوح

Hear the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation

Hanmo Chen, Chengcheng Liu, Tianxiao Chen وآخرون · 2026

Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sou …

نسخة أولية وصول مفتوح

What Makes an Efficient VLA? Navigating Action-Head Design, Scaling, and Latency

Luoyang Sun, Guoyang Xia, Fengfa Li وآخرون · 2026

Vision-Language-Action (VLA) models combine a pretrained vision encoder, a language backbone, and an action head, but their relative contribution has not been established under controlled, latency-paired conditions. We fix the backbone families (SigLIP2 and Qwen2.5) and the training pipeline, sweep action-head design a …

المؤلفون المشاركون