الباحثون

Zhiyuan Wang

المنشورات 2

نسخة أولية وصول مفتوح

VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

Yu Huang, Jungang Li, Zhiyuan Wang وآخرون · 2026

Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an …

نسخة أولية وصول مفتوح

EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action

Hao Wang, Jiajun Wen, Jingzhi Liu وآخرون · 2026

Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert. We present EWAM, an action-centric unified embo …

المؤلفون المشاركون