نسخة أولية وصول مفتوح
Learning Multimodal Embeddings with Evidence-Aligned Readout
Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we …