نسخة أولية وصول مفتوح
Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training
Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, diluting learning signals for informativ …