Abstract

Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational overhead. Consequently, their practical production is hindered by generation latency in the cloud deployment like NVIDIA-GB200, alongside strict memory limits that pose further challenges at the edge device like DGX-Spark. To address these diverse hardware bottlenecks from cloud to edge device, we present a full-stack inference pipeline that integrates efficient algorithmic design with optimized operator implementations. Algorithmically, we introduce a cross-resolution two-stage generation scheduler that exploits the step-wise nature of diffusion: early low-resolution steps rapidly establish the global layout, while later high-resolution steps focus refinements of local and perceptual details. These stages are connected by a learned latent-to-latent mapping module, completely eliminating the computationally expensive VAE decode-reencode cycle for resolution transferring cross different resolutions. For operator implementation, we deploy a Recursive Self-Improvement (RSI) loop that searches kernel fusions and memory layouts, evaluating latency together with numerical agreement. Together, these optimizations deliver up to 30x end-to-end speedup and 20% lower memory: a 5-second 1344x768 video with audio is generated 3.5x faster than real time on an 8xGB200 node, and in under a minute fully memory-resident on a single DGX Spark.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, Y., Yu, J., Chen, J., Li, H., Xue, S., Liu, H., Luo, P., Han, S., & Xie, E. (2026). Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge. https://omanscience.com/en/articles/sol-h3-recursive-self-improvement-for-minimax-h3-inference-acceleration-on-sol-engine-across-cloud-and-edge

MLA 9

Li, Yitong, et al. "Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge." https://omanscience.com/en/articles/sol-h3-recursive-self-improvement-for-minimax-h3-inference-acceleration-on-sol-engine-across-cloud-and-edge.

Chicago (author–date)

Li, Yitong, Jincheng Yu, Junsong Chen, Haopeng Li, Shuchen Xue, Haozhe Liu, Ping Luo, Song Han, and Enze Xie. 2026. "Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge." https://omanscience.com/en/articles/sol-h3-recursive-self-improvement-for-minimax-h3-inference-acceleration-on-sol-engine-across-cloud-and-edge.

Harvard

Li, Y., Yu, J., Chen, J., Li, H., Xue, S., Liu, H., Luo, P., Han, S. and Xie, E. (2026) 'Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge', Available at: https://omanscience.com/en/articles/sol-h3-recursive-self-improvement-for-minimax-h3-inference-acceleration-on-sol-engine-across-cloud-and-edge.

Vancouver

Li Y, Yu J, Chen J, Li H, Xue S, Liu H, et al. Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge. https://omanscience.com/en/articles/sol-h3-recursive-self-improvement-for-minimax-h3-inference-acceleration-on-sol-engine-across-cloud-and-edge

IEEE

Y. Li, J. Yu, J. Chen, H. Li, S. Xue, H. Liu, P. Luo, S. Han, and E. Xie, "Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge," https://omanscience.com/en/articles/sol-h3-recursive-self-improvement-for-minimax-h3-inference-acceleration-on-sol-engine-across-cloud-and-edge.