Preprint Open access
Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps
LLMs are released at a rapid pace, raising a natural question: how do two independently trained models relate, both in which layers correspond and in how features transform between them? We study this by learning an activation alignment, a map from a source model's layerwise activations to a target's. Our method, MATCH …