Search papers, labs, and topics across Lattice.
This paper introduces the Cross-confounder Robustness Margin (CRoMa), a novel measure for evaluating the robustness of pathology foundation models against systematic non-biological variations across clinical settings. Unlike the existing Robustness Index (RI), which relies on pooled data and overlooks sample-level heterogeneity, CRoMa provides a distributional perspective by comparing distances to biologically relevant matches and distractors at the sample level. The findings indicate that models exhibit significant variability in robustness profiles, highlighting the importance of considering lower-tail performance when selecting models for clinical deployment.
CRoMa reveals that model robustness is not just a single score but a nuanced distribution that highlights vulnerabilities to shortcut learning in pathology models.
Pathology foundation models are approaching clinical deployment, yet remain vulnerable to systematic non-biological variation across centres. Differences in tissue preparation, staining and scanning are strongly encoded in their representations, enabling shortcut learning and weakening generalisation across cohorts and institutions. The Robustness Index (RI) quantifies whether local representation geometry is dominated by biology or by non-biological variation, but its count-based formulation discards distance information. We show that adding distance weights changes little because the deeper limitation lies in RI's pooled, fixed-neighbourhood design, which obscures sample-level heterogeneity and effectively evaluates only a model-dependent subset of samples. We introduce the Cross-confounder Robustness Margin (CRoMa), a sample-resolved measure that directly compares distances to cross-confounder biological matches and same-confounder biological distractors. CRoMa recasts robustness as a cohort-wide margin distribution rather than a single pooled score. We evaluated frozen representations from 20 tile-level encoders across three benchmarks and 4 slide-level encoders on a fourth. Rankings by median CRoMa were broadly consistent across datasets, while the underlying distributions revealed substantial within-model heterogeneity. Every tile encoder retained a confounder-dominated lower tail, whose prevalence and severity varied markedly across models. These distinct robustness profiles frame model selection as a Pareto trade-off between typical and lower-tail robustness. Higher CRoMa was also associated with smaller shortcut-induced performance drops after supervised adaptation. By turning representation geometry into a distributional robustness readout that anticipates downstream shortcut susceptibility, CRoMa provides a principled basis for robustness assessment and model selection.