الباحثون

Song Han

المنشورات 9

نسخة أولية وصول مفتوح

Long-WAM: Scaling the Context of World-Action Models

Wei Huang, Bohan Zhang, Chenzhi Liu وآخرون · 2026

Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is …

نسخة أولية وصول مفتوح

Humanize: Judgement Engineering for Agentic Coding

Sihao Liu, Ligeng Zhu, Zijian Zhang وآخرون · 2026

Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done. We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the b …

نسخة أولية وصول مفتوح

Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model

Liming Lu, Xianzheng Ma, Wenkun He وآخرون · 2026

Video world models are expected to predict how the physical world evolves, yet they often produce visually plausible videos that violate basic physical principles. Existing approaches commonly assume that natural language is insufficient to represent the physical knowledge required for reliable generation, and therefor …

نسخة أولية وصول مفتوح

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang وآخرون · 2026

Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Pl …

نسخة أولية وصول مفتوح

Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs

Haozhan Tang, Hao Kang, Han Cai وآخرون · 2026

Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases an …

نسخة أولية وصول مفتوح

Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge

Yitong Li, Jincheng Yu, Junsong Chen وآخرون · 2026

Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational ov …

المؤلفون المشاركون