الباحثون

Xiaoshuang Shi

المنشورات 2

نسخة أولية وصول مفتوح

Dynamic Alignment and Calibration for Multimodal Learning

Jinghao Xu, Zhenhua Guo, Xiaofeng Zhu وآخرون · 2026

Dynamic multimodal learning aims to learn robust representations by adaptively modeling information discrepancies across modalities. However, existing methods still suffer from two limitations: (i) static cross-modal alignment strategies usually impose uniform constraints on all samples while overlooking sample-wise va …

نسخة أولية وصول مفتوح

VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

Yu Huang, Jungang Li, Zhiyuan Wang وآخرون · 2026

Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an …

المؤلفون المشاركون