Preprint Open access
Training-Free Transformer Merging via Sequential Local Operator Alignment
Training-free model merging aims to combine multiple fine-tuned models into a single model without further optimization on labeled data. Yet, in transformers, independently merging individual layers can affect a shared attention computation because the query-key and value-output operators depend on composed matrices, o …