Abstract

Mixture-of-Experts (MoE) architectures allow frontier language models to scale to trillions of parameters, but their deployment is constrained by massive memory footprints and memory-bandwidth limitations. Although modern accelerators provide Sparse Tensor Cores (SpTCs) that reduce weight storage and increase throughput through low-precision semi-structured sparsity, exploiting them for MoEs remains challenging because of substantial model-quality degradation and the lack of grouped sparse GEMM primitives. We present an end-to-end hardware-software co-design framework that compresses expert weights into hardware-native, low-precision sparse representations and accelerates their execution on SpTCs. Algorithmically, our framework relaxes discrete semi-structured support selection through continuous reparameterization, enabling differentiable joint optimization with quantized weights under a router-weighted reconstruction objective and scalable expert-parallel compression. Systemically, we develop a custom grouped sparse GEMM kernel tailored to low-precision sparse MoE inference on SpTCs. Across MoE models ranging from 30 billion to one trillion parameters, our framework improves state-of-the-art joint sparse-quantization accuracy by up to 4.35 percentage points while preserving 96.09% of the original model's performance. On NVIDIA B200 GPUs, our kernel outperforms the vendor baseline by up to $1.65\times$, increasing serving throughput by $1.18\times$ and reducing end-to-end latency by up to $4.03\times$. These results establish hardware-software co-design as a practical path toward scalable and efficient MoE deployment.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Lee, K., Lee, N., & Alistarh, D. (2026). Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts. https://omanscience.com/en/articles/hardware-native-joint-sparse-quantization-for-trillion-scale-mixture-of-experts

MLA 9

Lee, Kwanhee, et al. "Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts." https://omanscience.com/en/articles/hardware-native-joint-sparse-quantization-for-trillion-scale-mixture-of-experts.

Chicago (author–date)

Lee, Kwanhee, Namhoon Lee, and Dan Alistarh. 2026. "Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts." https://omanscience.com/en/articles/hardware-native-joint-sparse-quantization-for-trillion-scale-mixture-of-experts.

Harvard

Lee, K., Lee, N. and Alistarh, D. (2026) 'Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts', Available at: https://omanscience.com/en/articles/hardware-native-joint-sparse-quantization-for-trillion-scale-mixture-of-experts.

Vancouver

Lee K, Lee N, Alistarh D. Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts. https://omanscience.com/en/articles/hardware-native-joint-sparse-quantization-for-trillion-scale-mixture-of-experts

IEEE

K. Lee, N. Lee, and D. Alistarh, "Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts," https://omanscience.com/en/articles/hardware-native-joint-sparse-quantization-for-trillion-scale-mixture-of-experts.