Search papers, labs, and topics across Lattice.
This paper introduces CORAM, a novel method for merging finetuned models that enhances the integration of specialized capabilities without requiring joint training or access to original data. By utilizing a unique approach that partitions target matrices into row slices and employs singular value decomposition within the base-model framework, CORAM effectively addresses the limitations of existing orthogonal merging techniques. The method demonstrates significant improvements over OrthoMerge, achieving performance gains of 0.25 to 1.35 points across various model families and tasks, while also matching or exceeding traditional weight-space baselines.
CORAM achieves up to 1.35 points of performance improvement in model merging, redefining the boundaries of orthogonal transformations in AI.
Merging finetuned models combines specialized capabilities without joint training or access to the original data. Most methods operate by linear arithmetic in Euclidean weight space, which cannot carry the geometry of the update. Orthogonal Model Merging (OrthoMerge) uses a single orthogonal transform for each weight matrix, but such a transform cannot change singular values. We propose CORAM, which partitions each target matrix into row slices, represents every expert slice by its singular value decomposition in the corresponding base-model SVD frame, and merges the task-specific factors on their corresponding manifolds. Because manifold averaging contracts the merged update, CORAM applies an amplification coefficient $位=魏\hat{c}$. The scale c_hat is estimated from the expert and merged update norms and is approximately $\sqrt{N}$ for $N$ experts with comparable update magnitudes. The restoration strength kappa is selected from the dispersion of expert updates without evaluating candidate merged models. This rule remains within 0.72 points of the best swept value on all evaluated suites. CORAM also includes spread slicing to distribute highly updated rows across slices and a residual pathway for non-target layers. Across four suites covering three model families, 3B to 9B scales, and language and vision-language experts, CORAM improves over OrthoMerge by 0.25 to 1.35 points and matches or exceeds the strongest weight-space baselines.