الملخص

LLMs are released at a rapid pace, raising a natural question: how do two independently trained models relate, both in which layers correspond and in how features transform between them? We study this by learning an activation alignment, a map from a source model's layerwise activations to a target's. Our method, MATCHA, factors this map into a layer map, whose output is an explicit target-by-source matrix that can be extracted and inspected, and a layer-shared feature map between hidden spaces. Most of prior work fixes the layer correspondence in advance, pairing layers at roughly the same relative depth; in contrast, we learn both factors jointly from prompts. Across 42 pairs of seven models spanning three different families, MATCHA reconstructs the target's activations more faithfully and improves retrieval-based metrics substantially, w.r.t. previous approaches. The recovered maps are broadly monotone in depth but, in contrast with most previous approaches, are consistently many-to-many: each target layer draws on a band of source layers. Our alignments also enable transfer of activation-space interventions, allowing steering vectors and probes developed for one model to transfer to another.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Sudakov, A., Bar-Shalom, G., Frasca, F., & Maron, H. (2026). Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps. https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps

MLA 9

Sudakov, Alina, et al. "Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps." https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps.

شيكاغو (المؤلف–التاريخ)

Sudakov, Alina, Guy Bar-Shalom, Fabrizio Frasca, and Haggai Maron. 2026. "Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps." https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps.

هارفارد

Sudakov, A., Bar-Shalom, G., Frasca, F. and Maron, H. (2026) 'Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps', Available at: https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps.

فانكوفر

Sudakov A, Bar-Shalom G, Frasca F, Maron H. Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps. https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps

IEEE

A. Sudakov, G. Bar-Shalom, F. Frasca, and H. Maron, "Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps," https://omanscience.com/ar/articles/learning-cross-model-activation-alignments-with-explicit-many-to-many-layer-maps.