A preprint posted to arXiv on September 29 proposes SMAT, a merge-aware training method that aims to boost the performance of merged expert models while keeping training overhead below 2% of standard fine-tuning. arXiv
The technique addresses a known limitation in model merging: individually fine-tuned experts often underperform after merging because their training does not anticipate the operations that combine them. Model merging integrates multiple specialized models without joint retraining, using operations such as scaling a model’s update, masking coordinates, or adding perturbations from other experts. Standard fine-tuning optimizes only for the individual task, not for merged performance. Existing merge-aware training (MAT) methods have not fully captured these common merging operations or have added meaningful training cost.
Automatica reported this week that Allen AI released its Olmo-core 3 training stack for mixture-of-experts models, which tackles multi-expert scaling from the training side, rather than through post-hoc merging. SMAT, in contrast, focuses on making the experts themselves more mergeable.
The SMAT method jointly optimizes the expert’s task loss and the expected loss at simulated merged parameters. At each training step, it generates these simulated parameters by sampling scaling coefficients, masks, and additive noise. Efficiency techniques—periodic scheduling, kernel fusion, and parameter storage switching—keep the computational cost to one forward and one backward pass per step, according to the preprint.
Across four language and vision-language backbones, SMAT improved the mean score across five merging methods by 1.07 to 2.16 points over the strongest baseline for each backbone, the authors report. The training-time overhead was less than 2% over standard fine-tuning.
The preprint has not been peer-reviewed. The reported gains are based on the authors’ own experiments; no independent evaluations are provided. A code repository is linked in the paper, but its contents were not independently verified.