الباحثون

Yang Cao

المنشورات 6

نسخة أولية وصول مفتوح

Towards Unified Evaluation of Prompt Enhancers for Video Generation

Yawen Shao, Yubo Zhu, Ziyun Dai وآخرون · 2026

Modern video generators can realize increasingly complex visual narratives, positioning the prompt enhancer (PE) as a critical bridge from concise user instructions and multimodal references to structured cinematic plans. However, existing PE evaluation relies on rendered videos, imposing substantial computational and …

نسخة أولية وصول مفتوح

SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning

Yang Cao, Jiaxin Zhang, Dave Zhenyu Chen وآخرون · 2026

Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more …

نسخة أولية وصول مفتوح

PhoenixSR: Generative Heterogeneous Distillation Unleashes Efficient Models for Real-World Super-Resolution

Xin Di, Mingyu Shi, Yuanfei Bao وآخرون · 2026

Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while preserving faithful content. Diffusion-based SR benefits from strong generative priors but incurs substantial computational overhead, whereas feed-forward CNN and Trans …

نسخة أولية وصول مفتوح

Code Plans, Diffusion Renders: Open-Ended Generative World Modeling

Zixun Fang, Yawen Shao, Kai Zhu وآخرون · 2026

We introduce \textbf{CoDeR}, a new paradigm for world modeling. Unlike existing video world models that implicitly represent world dynamics through visual observations, our system explicitly constructs an executable world with code and employs video generation models for visual realization. Specifically, we coordinate …

نسخة أولية وصول مفتوح

The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts

Manjiang Yu, Hongji Li, Zihan Wang وآخرون · 2026

Linear probing and activation steering use linear directions to predict and control behavior-level concepts such as correctness, safety, and social bias in question answering. We call these \emph{behavior-level concept directions}. Drawing on the Linear Representation Hypothesis (LRH) and circuit studies, these directi …

المؤلفون المشاركون