Preprint Open access
Conditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional Teacher
Causal video diffusion models generate video autoregressively, which suits streaming, interactive, and long-video generation. Under standard training, however, they often yield lower generation quality than bidirectional models of the same size. Many existing approaches address this gap by initializing from or distilli …