Abstract
Unified language models are increasingly expected to combine heterogeneous capabilities, such as mathematics, code, instruction following, and controllable thinking behavior, within a single set of parameters. A common solution is sequential post-training on multiple objectives, but this entangles all objectives along one optimization trajectory and makes the final model highly sensitive to training order, data ratios, schedules, and stopping criteria. Weight-space merging offers a modular alternative, but naive merging of single-objective experts often fails: domain capabilities degrade sharply, or think/non-think modes collapse into one dominant behavior. We attribute both failures to incompatible weight-space geometry: experts trained on single objectives drift to distant regions of parameter space, placing their interpolations outside any shared low-loss basin. We propose Mixture-Trained Merging (MTM), which trains each branch on an objective-biased data mixture rather than a single objective, exposing it to cross-objective interactions and making branches compatible at merge time. MTM uses merged-model evaluations as a low-cost signal for selecting branch mixtures, avoiding expensive data-mixture ablations. The procedure is iterative: each round promotes the base model using globally selected merge coefficients and refines each branch mixture using domain-preferred coefficients under constraints that preserve other objectives. To scale beyond simplex grid search, MTM uses qNEHVI-based multi-objective Bayesian optimization. Across code, mathematics, instruction following, and think/non-think control, MTM outperforms naive merging and preserves behavioral separation where single-objective merging collapses, suggesting that effective unified models require branches trained to be mergeable.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Kim, S., Jang, C., Lee, S., Ham, J., Bak, Y., Kim, B., & Lee, J. (2026). Mixture-Trained Merging for Unified Multi-Objective Models. https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models
MLA 9
Kim, SeongHyeon, et al. "Mixture-Trained Merging for Unified Multi-Objective Models." https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models.
Chicago (author–date)
Kim, SeongHyeon, Chaeyun Jang, Seungyoo Lee, Jiyeon Ham, Yunju Bak, Boseop Kim, and Juho Lee. 2026. "Mixture-Trained Merging for Unified Multi-Objective Models." https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models.
Harvard
Kim, S., Jang, C., Lee, S., Ham, J., Bak, Y., Kim, B. and Lee, J. (2026) 'Mixture-Trained Merging for Unified Multi-Objective Models', Available at: https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models.
Vancouver
Kim S, Jang C, Lee S, Ham J, Bak Y, Kim B, et al. Mixture-Trained Merging for Unified Multi-Objective Models. https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models
IEEE
S. Kim, C. Jang, S. Lee, J. Ham, Y. Bak, B. Kim, and J. Lee, "Mixture-Trained Merging for Unified Multi-Objective Models," https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models.