الملخص
LLMs are released at a rapid pace, raising a natural question: how do two independently trained models relate, both in which layers correspond and in how features transform between them? We study this by learning an activation alignment, a map from a source model's layerwise activations to a target's. Our method, MATCHA, factors this map into a layer map, whose output is an explicit target-by-source matrix that can be extracted and inspected, and a layer-shared feature map between hidden spaces. Most of prior work fixes the layer correspondence in advance, pairing layers at roughly the same relative depth; in contrast, we learn both factors jointly from prompts. Across 42 pairs of seven models spanning three different families, MATCHA reconstructs the target's activations more faithfully and improves retrieval-based metrics substantially, w.r.t. previous approaches. The recovered maps are broadly monotone in depth but, in contrast with most previous approaches, are consistently many-to-many: each target layer draws on a band of source layers. Our alignments also enable transfer of activation-space interventions, allowing steering vectors and probes developed for one model to transfer to another.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Sudakov, A., Bar-Shalom, G., Frasca, F., & Maron, H. (2026). Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps. https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps
MLA 9
Sudakov, Alina, et al. "Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps." https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps.
شيكاغو (المؤلف–التاريخ)
Sudakov, Alina, Guy Bar-Shalom, Fabrizio Frasca, and Haggai Maron. 2026. "Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps." https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps.
هارفارد
Sudakov, A., Bar-Shalom, G., Frasca, F. and Maron, H. (2026) 'Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps', Available at: https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps.
فانكوفر
Sudakov A, Bar-Shalom G, Frasca F, Maron H. Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps. https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps
IEEE
A. Sudakov, G. Bar-Shalom, F. Frasca, and H. Maron, "Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps," https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps.