الباحثون

Tianfan Xue

المنشورات 7

نسخة أولية وصول مفتوح

VDOT++: Unified Few-Step Video Generation via Unbalanced Optimal Transport Distillation

Yutong Wang, Xingtong Ge, Enhuai Liu وآخرون · 2026

Video creation spans text-to-video (T2V), image-to-video (I2V), and condition-based generation, yet video diffusion models remain costly because they repeatedly evaluate large backbones during sampling. Distribution matching distillation (DMD) reduces this cost, but its reverse Kullback--Leibler (KL) objective can prov …

نسخة أولية وصول مفتوح

DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation

Jiahao Zhan, Yan Wang, Yongrui Ma وآخرون · 2026

Streaming video generation has benefited from distribution matching distillation (DMD), which matches the joint distribution of video frames to a video teacher's approximation of the real video distribution. Although this joint matching mitigates drift during autoregressive rollouts, limitations remain in visual qualit …

نسخة أولية وصول مفتوح

Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation

Chenjian Gao, Zhihao Hu, Jianqi Ma وآخرون · 2026

Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rol …

نسخة أولية وصول مفتوح

PulseQuant: Propagation-Guided Subspace Correction for 4-Bit Video Diffusion Transformers

Yutong Wang, Xingtong Ge, Enhuai Liu وآخرون · 2026

Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction error an incomplete predictor of final impact. We introduce PulseQuant, a 4-bit post-training quantization method that combines trajectory sensitivity with activation geometry …

نسخة أولية وصول مفتوح

InternW0-$Δ$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

Xingyu Miao, Zizun Li, Baole Fang وآخرون · 2026

World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We in …

نسخة أولية وصول مفتوح

InternW0: A Foundational Physical World Model for Efficient Real-World Interactions

Jisong Cai, Yao Mu, Ganlin Yang وآخرون · 2026

Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi- …

المؤلفون المشاركون