الملخص
Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher. This self-supervised signal encourages the student to extract semantic information and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation. At 480p, Salt++ improves visual and motion quality by 57% and 45% over OmniForcing on JavisBench under the same 4-step causal setting. A separate scale-wise post-training stage extends Salt++ to 4-step $1664\times960$ generation, outperforming bidirectional LTX-2 on six of seven reported metrics. Project page: https://xingtongge.github.io/Saltpp
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Ge, X., Wang, Y., Zhu, L., Lin, H., Lin, F., Huang, Y., Zhang, X., Zhang, Y., Liu, Y., & Zhang, J. (2026). Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation. https://omanscience.com/ar/articles/salt-context-aligned-post-training-for-few-step-streaming-multimodal-generation
MLA 9
Ge, Xingtong, et al. "Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation." https://omanscience.com/ar/articles/salt-context-aligned-post-training-for-few-step-streaming-multimodal-generation.
شيكاغو (المؤلف–التاريخ)
Ge, Xingtong, Yutong Wang, Lunjie Zhu, Haitao Lin, Fangyu Lin, Yushi Huang, Xin Zhang, Yi Zhang, Yu Liu, and Jun Zhang. 2026. "Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation." https://omanscience.com/ar/articles/salt-context-aligned-post-training-for-few-step-streaming-multimodal-generation.
هارفارد
Ge, X., Wang, Y., Zhu, L., Lin, H., Lin, F., Huang, Y., Zhang, X., Zhang, Y., Liu, Y. and Zhang, J. (2026) 'Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation', Available at: https://omanscience.com/ar/articles/salt-context-aligned-post-training-for-few-step-streaming-multimodal-generation.
فانكوفر
Ge X, Wang Y, Zhu L, Lin H, Lin F, Huang Y, et al. Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation. https://omanscience.com/ar/articles/salt-context-aligned-post-training-for-few-step-streaming-multimodal-generation
IEEE
X. Ge, Y. Wang, L. Zhu, H. Lin, F. Lin, Y. Huang, X. Zhang, Y. Zhang, Y. Liu, and J. Zhang, "Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation," https://omanscience.com/ar/articles/salt-context-aligned-post-training-for-few-step-streaming-multimodal-generation.