[
    {
        "id": "osp-19859",
        "type": "article-journal",
        "title": "ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models",
        "author": [
            {
                "family": "Liu",
                "given": "Lianjun"
            },
            {
                "family": "Li",
                "given": "Shipeng"
            },
            {
                "family": "Huang",
                "given": "You"
            },
            {
                "family": "Yan",
                "given": "Weiqi"
            },
            {
                "family": "Qiu",
                "given": "Mingte"
            },
            {
                "family": "Liu",
                "given": "Huazhong"
            },
            {
                "family": "Zhu",
                "given": "Xiaofeng"
            },
            {
                "family": "Zhong",
                "given": "Yunshan"
            }
        ],
        "URL": "https://omanscience.com/en/articles/itc-moe-importance-guided-token-aware-compression-for-moe-diffusion-language-models",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Mixture-of-Experts (MoE) Diffusion Language Models (DLMs) offer flexible parallel decoding and increased model capacity, but their large number of expert parameters incurs substantial computation and storage costs. Existing low-rank MoE compression methods largely rely on static factorization and fixed rank allocation, which overlook the distinctive properties of MoE DLMs. Specifically, we identify two properties: cross-mode non-uniform redundancy, where parameter redundancy and sensitivity to rank truncation vary across the input, output, and expert modes, and token-wise utilization variation, where hot and cold tokens exhibit distinct spectral characteristics and expert activation patterns. To address these challenges, we propose ITC-MoE, an Importance-guided Token-aware Compression framework for MoE DLMs. ITC-MoE consists of two complementary components. First, Importance-guided Adaptive Tucker Compression (IATC) incorporates activation and gradient importance into expert weight transformation, jointly factorizes expert weights across multiple modes, and adaptively allocates ranks under a fixed parameter budget. Second, Token-aware Compensation and Routing (TCR) applies lightweight low-rank compensation to compression-sensitive hot tokens and restricts the candidate expert set for cold tokens with concentrated routing patterns. By jointly adapting compression capacity and inference execution to both parameter redundancy and token-wise variation, ITC-MoE substantially reduces the computation and storage costs of MoE DLMs while preserving their generation quality. For example, on SDAR-30B-A3B-Chat-b32, ITC-MoE maintains an accuracy of 96.33% on MultiArith under a 30% compression budget, while achieving up to a 7.22x end-to-end speedup. The code is publicly available at https://github.com/lianjunl13-sudo/ITC-MoE."
    }
]