Search papers, labs, and topics across Lattice.
This paper introduces EpiMer, a model merging framework that formulates the merging process as solving for the Fréchet mean on a Riemannian manifold, using the expected Hessian as the metric to capture the loss landscape geometry. By restricting computation to a low-rank subspace spanned by task vectors, EpiMer connects local curvature to epistemic uncertainty and provides theoretical error bounds decomposing merging error into subspace Fréchet variance and residual energy. Empirically, EpiMer outperforms existing methods when merging fine-tuned CLIP-ViT models across various image classification tasks.
Curvature-aware model merging, previously intractable, can now be efficiently approximated in a low-rank subspace, provably outperforming flat-geometry methods and unifying prior spectral techniques.
Model merging offers a promising avenue for knowledge integration and parallel development without retraining. Yet, existing methods either ignore the geometry of the loss landscape or rely on intractable full-space Hessian approximations. We propose EpiMer, a framework that casts model merging as solving the Fréchet mean on a Riemannian manifold and restricts the computation to a low-rank subspace spanned by the task vectors. With the expected Hessian as the metric, we reveal a connection between local curvature and epistemic uncertainty of the parameters. Our theoretical analysis decomposes the merging error bound into the subspace Fréchet variance and the residual energy, and provides a closed-form characterization of when curvature-aware merging provably outperforms flat-geometry methods. In addition, our framework unifies both curvature-aware methods and recent spectral methods as special cases of the subspace Fréchet mean with different geometric metrics. Merging fine-tuned CLIP-ViT models on eight image classification tasks, Epistemic Merging strictly outperforms the baselines on all three CLIP-ViT backbones at matched rank, improving the across-task average accuracy and worst-task accuracy on every backbone.