[
    {
        "id": "osp-21855",
        "type": "article-journal",
        "title": "Routing in Gradient Space: Balanced Usage Is Not Expert Specialization",
        "author": [
            {
                "family": "Li",
                "given": "Yuchen"
            },
            {
                "family": "Du",
                "given": "Mingyu"
            },
            {
                "family": "Fan",
                "given": "Zongqi"
            },
            {
                "family": "Tran",
                "given": "Nguyen H."
            },
            {
                "family": "Yong",
                "given": "Ken-Tye"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/routing-in-gradient-space-balanced-usage-is-not-expert-specialization",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose load-normalized router objective rewards grouping observations with aligned gradients. On five multi-task text-classification mixtures, we compare GAR with task-loss-only routing, gradient-combination and gradient-conflict methods, and load-balancing losses. With a fully trainable RoBERTa backbone and classification-head experts, GAR has the highest aggregate validation accuracy, 1.07 percentage points above task-loss-only routing. With frozen DeBERTa and Qwen3-1.7B backbones and low-rank adapter experts, it again ranks first, 1.10 points above task-loss-only routing, with better-balanced expert load and higher gradient-mass purity, the share of each expert's gradient-norm mass from its dominant task; the load-balancing losses flatten load further but leave this purity near its task-loss-only level. Top-1 routing, trainable full-parameter feed-forward network (FFN) experts, and a larger backbone also show positive aggregate gains. The results distinguish expert-load balance from gradient-based routing organization and indicate the predictive value of gradient-informed routing in multi-task text classification."
    }
]