[
    {
        "id": "osp-16180",
        "type": "article-journal",
        "title": "Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness",
        "author": [
            {
                "family": "Sui",
                "given": "Songyuan"
            },
            {
                "family": "Tan",
                "given": "Zhen"
            },
            {
                "family": "Zhang",
                "given": "Mohan"
            },
            {
                "family": "Khan",
                "given": "Rana Muhammad Shahroz"
            },
            {
                "family": "Hu",
                "given": "Xia"
            },
            {
                "family": "Chen",
                "given": "Tianlong"
            }
        ],
        "URL": "https://omanscience.com/en/articles/do-more-modalities-always-help-a-geometric-perspective-on-missing-modality-robustness",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Missing modality remains a longstanding challenge in multimodal learning. Existing methods typically address this issue through modality recovery or adaptive strategies. However, they overlook models' internal cross-modal dependencies formed during multimodal training, which later impair robustness. We systematically characterize a counterintuitive deployment-time failure mode: models trained on full modalities can underperform unimodal models when one modality is missing at inference time. This pattern appears across diverse architectures, such as fusion models, CLIP-style two-tower models, and vision-language models. We show that such degradation is closely associated with learned cross-modal dependencies in the principal parameter subspaces. Multimodal training induces structured rotations of these subspaces, particularly in cross-modal interaction layers. These rotations are associated with reduced task-aligned margins and larger task-aware representation harm under missing-modality inputs. We propose Geodesic Unlearning (GU), a lightweight parameter-editing method that leverages Grassmannian subspace geometry for structured subspace correction to improve missing-modality robustness. It rotates the principal input subspace toward a unimodal reference along a geodesic path. We prove that this correction minimizes the distance to the reference within a fixed subspace-distance budget. Experiments across architectures and datasets show that GU improves performance under missing-modality inference while preserving full-modality accuracy, outperforming strong missing-modality robustness baselines. These findings support a geometric view of deployment-time missing-modality degradation and suggest localized subspace editing as a practical route for robustness correction."
    }
]