نسخة أولية وصول مفتوح
Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Me …