الباحثون

Zhongang Cai

المنشورات 3

نسخة أولية وصول مفتوح

Looped Diffusion Transformer

Yong Xien Chng, Tianyi Chen, Wenwen Tong وآخرون · 2026

Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the par …

نسخة أولية وصول مفتوح

One-Step Next-Latent Prediction Is Not a World Model

Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a cond …

نسخة أولية وصول مفتوح

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Yijia Fan, Ziqi Huang, Zhongang Cai وآخرون · 2026

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be …

المؤلفون المشاركون