Abstract
Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation geometry. But then, do we even need paired examples for cross-modal alignment? Remarkably, we show that paired examples are unnecessary for coarse cross-modal alignment. Our simple Wasserstein Procrustes method with a coarse geometric initialization aligns two disjoint embedding sets by estimating a single orthogonal map without seeing any pairs. Across datasets, modalities, and unimodal models, we show that we can consistently align independently trained representations without pairs, and standard geometric alignment metrics accurately predict when this is possible. Nevertheless, we can naturally benefit from paired examples. In the very few-pair regime, our method substantially outperforms existing ones, while staying competitive with pair-based methods with more added examples. Finally, we demonstrate that the resulting alignments can enable text-to-image generation without paired examples. These results show that independently trained models often share enough geometry to establish cross-modal correspondence with little or no paired data.
Keywords
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Schnaus, D., Dagès, T., Cremers, D., Wang, X., & Isola, P. (2026). Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data. https://omanscience.com/en/articles/shared-geometry-as-a-rosetta-stone-cross-modal-alignment-without-paired-data
MLA 9
Schnaus, Dominik, et al. "Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data." https://omanscience.com/en/articles/shared-geometry-as-a-rosetta-stone-cross-modal-alignment-without-paired-data.
Chicago (author–date)
Schnaus, Dominik, Thomas Dagès, Daniel Cremers, Xi Wang, and Phillip Isola. 2026. "Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data." https://omanscience.com/en/articles/shared-geometry-as-a-rosetta-stone-cross-modal-alignment-without-paired-data.
Harvard
Schnaus, D., Dagès, T., Cremers, D., Wang, X. and Isola, P. (2026) 'Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data', Available at: https://omanscience.com/en/articles/shared-geometry-as-a-rosetta-stone-cross-modal-alignment-without-paired-data.
Vancouver
Schnaus D, Dagès T, Cremers D, Wang X, Isola P. Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data. https://omanscience.com/en/articles/shared-geometry-as-a-rosetta-stone-cross-modal-alignment-without-paired-data
IEEE
D. Schnaus, T. Dagès, D. Cremers, X. Wang, and P. Isola, "Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data," https://omanscience.com/en/articles/shared-geometry-as-a-rosetta-stone-cross-modal-alignment-without-paired-data.