Search papers, labs, and topics across Lattice.
This study introduces the concept of generative-process diversity as a critical measure for understanding the correlated failures of language model pairs, contrasting it with traditional assessments based on semantic similarity. By employing Normalised Compression Distance to quantify this diversity across 38 language models, the authors reveal that higher generative-process diversity correlates with reduced chances of simultaneous failure across various benchmark tasks. The findings underscore that traditional measures of semantic diversity are insufficient for predicting model resilience, highlighting the importance of generative processes in multi-model systems.
Increased generative-process diversity in language models significantly reduces correlated failures, a factor overlooked by traditional semantic similarity assessments.
Diversity is a widely observed factor in the resilient function of collective systems, yet the type of diversity that matters depends on the properties and failure modes of the system. This distinction is important for systems composed of multiple language models. Different models may be treated as independent components even when their behaviour and failures remain strongly correlated. Assessments of language-model populations using semantic similarity demonstrate limited semantic diversity, but this captures only differences in the meaning of observed outputs. We argue that a more fundamental notion of model diversity is generative-process diversity, the differences between processes capable of generating the observed outputs. Drawing from Algorithmic Information Theory, we use Normalised Compression Distance between raw model outputs, residualised against a permutation control, as a measure of inferred generative-process diversity. Across 38 language models, this measure identifies population structure missed by semantic similarity and predicts cross-task variation in chance-corrected correlated failure among model pairs across ten disjoint benchmark families, beyond semantic similarity and model-pair capability. The cross-benchmark partial rank association is $-0.216$ with a 95% interval of $[-0.309,-0.122]$, and the estimate is negative on all ten benchmarks. These results indicate that increased generative-process diversity is associated with reduced correlated failure in model pairs that is not attributable to semantic similarity or capability. Inferred generative-process diversity offers a novel and practical approach for investigating diversity of multi-model systems in safety-relevant contexts.