الملخص

Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visual quality against motion, and that restoring gradients through the history aligns causal training far more closely with bidirectional training. Motivated by these observations, we introduce Self-Aligned Forcing (SAF), a training scheme that aligns the history of each block with the noise level of the block being denoised. Specifically, the history is the K/V produced by preceding blocks at the same denoising stage, so all blocks at a stage can be denoised in a single forward pass under a causal mask. This keeps the noisy history differentiable, allowing future losses to optimize how it is encoded. SAF therefore avoids a separate no-gradient rollout and per-block timestep-zero recaching, training up to 1.8x faster than prior methods with lower memory. At inference, SAF achieves the highest single-GPU throughput among existing methods and keeps one history bank per stage for a multi-GPU pipeline, reaching 49.1 FPS on 4 GPUs. Experiments show superior long-horizon generation with a better balance between visual quality and motion. Project page: https://anonymous.4open.science/w/self-aligned-forcing/.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Wang, W., Chen, Z., Dai, Y., Li, B., Zhang, Y., Rahmani, H., Ke, Q., & Cai, J. (2026). Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History. https://omanscience.com/ar/articles/self-aligned-forcing-streaming-video-diffusion-with-differentiable-noisy-history

MLA 9

Wang, Weiqiang, et al. "Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History." https://omanscience.com/ar/articles/self-aligned-forcing-streaming-video-diffusion-with-differentiable-noisy-history.

شيكاغو (المؤلف–التاريخ)

Wang, Weiqiang, Zhuokun Chen, Yusheng Dai, Boying Li, Yi Zhang, Hossein Rahmani, Qiuhong Ke, and Jianfei Cai. 2026. "Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History." https://omanscience.com/ar/articles/self-aligned-forcing-streaming-video-diffusion-with-differentiable-noisy-history.

هارفارد

Wang, W., Chen, Z., Dai, Y., Li, B., Zhang, Y., Rahmani, H., Ke, Q. and Cai, J. (2026) 'Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History', Available at: https://omanscience.com/ar/articles/self-aligned-forcing-streaming-video-diffusion-with-differentiable-noisy-history.

فانكوفر

Wang W, Chen Z, Dai Y, Li B, Zhang Y, Rahmani H, et al. Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History. https://omanscience.com/ar/articles/self-aligned-forcing-streaming-video-diffusion-with-differentiable-noisy-history

IEEE

W. Wang, Z. Chen, Y. Dai, B. Li, Y. Zhang, H. Rahmani, Q. Ke, and J. Cai, "Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History," https://omanscience.com/ar/articles/self-aligned-forcing-streaming-video-diffusion-with-differentiable-noisy-history.