Preprint Open access
Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data
Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation ge …