[
    {
        "id": "osp-23756",
        "type": "article-journal",
        "title": "Correlation-Guided Encoder Selection for Multi-Encoder Large Audio-Language Models",
        "author": [
            {
                "family": "Liao",
                "given": "Pei-Jun"
            },
            {
                "family": "Lee",
                "given": "Hung-Shin"
            },
            {
                "family": "Ren",
                "given": "Wenze"
            },
            {
                "family": "Hung",
                "given": "Kuo-Hsuan"
            },
            {
                "family": "Lee",
                "given": "Hung-yi"
            },
            {
                "family": "Wang",
                "given": "Hsin-Min"
            }
        ],
        "URL": "https://omanscience.com/en/articles/correlation-guided-encoder-selection-for-multi-encoder-large-audio-language-models",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Multi-encoder fusion extends Large Audio-Language Models (LALMs) beyond speech-centric recognition, but selecting encoders via intuition or exhaustive search often introduces redundant representations and inflates an already constrained compute budget. We propose CUES (Correlation-gUided Encoder Selection), a lightweight heuristic that estimates complementarity through task- and category-level Pearson correlations between encoders' performance profiles, scoring a candidate set from single-encoder evaluations alone--without fusion training during selection. Evaluated on the XARES-LLM benchmark with a frozen SmolLM2-135M backbone (LoRA-adapted) via five-fold cross-validation, CUES consistently identifies the same configuration per track from held-out development splits alone, without using test data for selection. For the broad Track~A suite, CUES selects a cross-family trio (Whisper-medium, mHuBERT-147, and Dasheng-base), achieving a 4.3% relative gain over Whisper-medium (0.771 vs. 0.739). For Track~B text generation, it re-anchors on a focused, speech-only pair (mHuBERT-147 and WavLM-base-plus) and actively abstains from adding a divergent encoder, outperforming mHuBERT-147 by 6.3% (0.589 vs. 0.554). Rather than a failure to scale, this divergence is consistent with a diversity--interference trade-off that CUES navigates per track from correlation signals alone: across the evaluated pool, added cross-family diversity tends toward an inverted-U on broad audio tasks but toward steady degradation on text generation, which favors a focused, speech-anchored set."
    }
]