Abstract
Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examining DMD's teacher and fake-score signals, we find that weak re-noising keeps the teacher posterior concentrated near motion-deficient rollouts, limiting motion-restoring guidance. Meanwhile, stronger-motion rollouts tend to incur larger fake-score fitting errors, which can hinder the generator's learning of interaction dynamics. We propose DyMD, a DMD framework that adapts both teacher supervision and critic fitting to the evolving student. Temporal affinity--conditioned re-noise sampling adapts the timestep distribution to each rollout's current interaction fidelity by mixing the base schedule with a teacher prior motivated by local posterior variation, thereby balancing motion recovery and appearance refinement. To better track stronger-motion rollouts, dynamics-guided fake-score tracking uses a noise-conditioned predictor to estimate noise-relative fitting difficulty from latent temporal dynamics, then upweights predicted-hard rollouts in the critic loss. Using DyMD, we distill a 14B teacher into a four-step 1.3B student with no auxiliary modules at inference. On embodied-video benchmarks, the student improves R-Bench task adherence by $9.6$ percentage points and PAI-Bench-G Domain score by $5.1$ points over Base DMD while maintaining comparable visual quality. As a backbone for downstream action planning, our student achieves 34% mean success across two WorldArena tasks, compared with 16% for Base DMD.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Xu, H., Huang, J., Lu, X., Zhong, M., Fan, Z., Huang, L., & Liu, S. (2026). DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models. https://omanscience.com/en/articles/dymd-preserving-interaction-dynamics-through-distribution-matching-distillation-in-few-step-video-world-models
MLA 9
Xu, Haojun, et al. "DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models." https://omanscience.com/en/articles/dymd-preserving-interaction-dynamics-through-distribution-matching-distillation-in-few-step-video-world-models.
Chicago (author–date)
Xu, Haojun, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan, Linjiang Huang, and Si Liu. 2026. "DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models." https://omanscience.com/en/articles/dymd-preserving-interaction-dynamics-through-distribution-matching-distillation-in-few-step-video-world-models.
Harvard
Xu, H., Huang, J., Lu, X., Zhong, M., Fan, Z., Huang, L. and Liu, S. (2026) 'DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models', Available at: https://omanscience.com/en/articles/dymd-preserving-interaction-dynamics-through-distribution-matching-distillation-in-few-step-video-world-models.
Vancouver
Xu H, Huang J, Lu X, Zhong M, Fan Z, Huang L, et al. DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models. https://omanscience.com/en/articles/dymd-preserving-interaction-dynamics-through-distribution-matching-distillation-in-few-step-video-world-models
IEEE
H. Xu, J. Huang, X. Lu, M. Zhong, Z. Fan, L. Huang, and S. Liu, "DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models," https://omanscience.com/en/articles/dymd-preserving-interaction-dynamics-through-distribution-matching-distillation-in-few-step-video-world-models.