[
    {
        "id": "osp-19834",
        "type": "article-journal",
        "title": "MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs",
        "author": [
            {
                "family": "Wang",
                "given": "Xudong"
            },
            {
                "family": "Wu",
                "given": "Hao"
            },
            {
                "family": "Hu",
                "given": "Haozhe"
            },
            {
                "family": "Yin",
                "given": "Peiran"
            },
            {
                "family": "Chen",
                "given": "Xinghao"
            },
            {
                "family": "Ma",
                "given": "Yunpu"
            },
            {
                "family": "Zhang",
                "given": "Wei"
            },
            {
                "family": "Shen",
                "given": "Xiaoyu"
            }
        ],
        "URL": "https://omanscience.com/en/articles/mwop-modality-aware-width-wise-operation-pruning-for-efficient-mllms",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a $1.6\\times$ prefill speedup with 99.7\\% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from $2.0\\times$ and $1.9\\times$ to $2.9\\times$ and $2.7\\times$, respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at https://github.com/EIT-NLP/MWOP."
    }
]