الباحثون

Xintong Han

المنشورات 2

نسخة أولية وصول مفتوح

Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training

Shuyuan Tu, Qi Tian, Yinming Huang وآخرون · 2026

Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, diluting learning signals for informativ …

نسخة أولية وصول مفتوح

Flow Matching Reinforcement for 3D Mesh Generation via Dynamic Homing Optimization

Zhen Zhou, Zhiwei Ning, Puhua Jiang وآخرون · 2026

Flow matching is central to 3D generation, yet in practice its reinforcement learning (RL) methods are largely adapted from 2D visual generation. Representative DPO-, GRPO-, and NFT-style objectives, when applied to negative trajectories, mainly steer predicted velocities away from the corresponding directions without …

المؤلفون المشاركون