[
    {
        "id": "osp-24930",
        "type": "article-journal",
        "title": "FlatClip: A Geometry-Aware Surface-Level Baseline for fMRI Representation Learning",
        "author": [
            {
                "family": "Wang",
                "given": "Mo"
            },
            {
                "family": "Ye",
                "given": "Wenhao"
            },
            {
                "family": "Ning",
                "given": "Zihan"
            },
            {
                "family": "Zuo",
                "given": "Jiayu"
            },
            {
                "family": "Xia",
                "given": "Junfeng"
            },
            {
                "family": "Wen",
                "given": "Hongkai"
            },
            {
                "family": "Liu",
                "given": "Quanying"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/flatclip-a-geometry-aware-surface-level-baseline-for-fmri-representation-learning",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models are efficient but coarse, whereas voxel-level models preserve fine-grained spatial structure but require specialized 3D/4D architectures and costly fMRI-specific pretraining. We ask how effectively an image-pretrained encoder can reuse the spatial organization of cortical activity. Motivated by evidence that macroscale brain activity is strongly constrained by brain geometry, we introduce FlatClip, a frozen-encoder surface-level baseline that renders cortical activity as geometry-aware flatmap sequences and reuses a frozen SigLIP2 image encoder with only a lightweight downstream probe. Across resting-state benchmarks, FlatClip serves as a competitive middle-ground representation, outperforming ROI-level baselines on HCP and ADNI tasks while remaining weaker on PPMI and below the strongest voxel-level models overall. On visual-fMRI decoding, restricting the input to visual or NSD-provided task-active cortex improves performance, highlighting the value of task-relevant cortical coverage. Spatial perturbation controls reduce the predictive performance of flatmap features under both retrained and fixed readouts, and anatomy-linked arrangements consistently outperform vertex permutations across three colormaps. Together, these results position surface-level flatmap sequences as a practical middle-ground baseline between ROI and voxel models, and support the utility of anatomy-linked spatial organization for reusing image-pretrained features. Code is available at https://github.com/OneMore1/FlatClip."
    }
]