Abstract

Unified language models are increasingly expected to combine heterogeneous capabilities, such as mathematics, code, instruction following, and controllable thinking behavior, within a single set of parameters. A common solution is sequential post-training on multiple objectives, but this entangles all objectives along one optimization trajectory and makes the final model highly sensitive to training order, data ratios, schedules, and stopping criteria. Weight-space merging offers a modular alternative, but naive merging of single-objective experts often fails: domain capabilities degrade sharply, or think/non-think modes collapse into one dominant behavior. We attribute both failures to incompatible weight-space geometry: experts trained on single objectives drift to distant regions of parameter space, placing their interpolations outside any shared low-loss basin. We propose Mixture-Trained Merging (MTM), which trains each branch on an objective-biased data mixture rather than a single objective, exposing it to cross-objective interactions and making branches compatible at merge time. MTM uses merged-model evaluations as a low-cost signal for selecting branch mixtures, avoiding expensive data-mixture ablations. The procedure is iterative: each round promotes the base model using globally selected merge coefficients and refines each branch mixture using domain-preferred coefficients under constraints that preserve other objectives. To scale beyond simplex grid search, MTM uses qNEHVI-based multi-objective Bayesian optimization. Across code, mathematics, instruction following, and think/non-think control, MTM outperforms naive merging and preserves behavioral separation where single-objective merging collapses, suggesting that effective unified models require branches trained to be mergeable.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Kim, S., Jang, C., Lee, S., Ham, J., Bak, Y., Kim, B., & Lee, J. (2026). Mixture-Trained Merging for Unified Multi-Objective Models. https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models

MLA 9

Kim, SeongHyeon, et al. "Mixture-Trained Merging for Unified Multi-Objective Models." https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models.

Chicago (author–date)

Kim, SeongHyeon, Chaeyun Jang, Seungyoo Lee, Jiyeon Ham, Yunju Bak, Boseop Kim, and Juho Lee. 2026. "Mixture-Trained Merging for Unified Multi-Objective Models." https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models.

Harvard

Kim, S., Jang, C., Lee, S., Ham, J., Bak, Y., Kim, B. and Lee, J. (2026) 'Mixture-Trained Merging for Unified Multi-Objective Models', Available at: https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models.

Vancouver

Kim S, Jang C, Lee S, Ham J, Bak Y, Kim B, et al. Mixture-Trained Merging for Unified Multi-Objective Models. https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models

IEEE

S. Kim, C. Jang, S. Lee, J. Ham, Y. Bak, B. Kim, and J. Lee, "Mixture-Trained Merging for Unified Multi-Objective Models," https://omanscience.com/en/articles/mixture-trained-merging-for-unified-multi-objective-models.