الباحثون

Haiwen Diao

المنشورات 4

نسخة أولية وصول مفتوح

Looped Diffusion Transformer

Yong Xien Chng, Tianyi Chen, Wenwen Tong وآخرون · 2026

Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the par …

نسخة أولية وصول مفتوح

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi وآخرون · 2026

Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that expli …

نسخة أولية وصول مفتوح

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Yijia Fan, Ziqi Huang, Zhongang Cai وآخرون · 2026

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be …

نسخة أولية وصول مفتوح

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

Zhongbo Zhang, Jiayi Jin, Yifan Wang وآخرون · 2026

Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedur …

المؤلفون المشاركون