Search papers, labs, and topics across Lattice.
This study audits five diversity metrics used in majority voting among LLM ensembles to determine their effectiveness in predicting performance gains. By controlling for model capability across 31,900 subsets of LLMs, the authors reveal that while latent complementarity is prevalent, simple voting only outperforms the best individual model in a small fraction of cases. The findings indicate that diversity measures are largely entangled with model capability, challenging the assumption that diversity directly correlates with improved ensemble performance.
Majority voting in LLM ensembles might be more about capability than diversity, with only 9.98% of subsets showing performance gains over the best model.
Majority voting over LLMs is widely assumed to benefit from diversity, and diversity measures are used to choose which models to combine. We ask whether five such measures track diversity or mainly re-express capability, auditing them as predictors of majority-vote gain over the best member across 31,900 subsets of 30 LLMs on MMLU-Pro (29 on TruthfulQA) under explicit capability controls. Three findings emerge. First, latent complementarity is ubiquitous: oracle gain is positive in 100% of subsets, yet simple voting beats the strongest member in only 9.98% of all canonical size-3 subsets (18.71% with held-out best selection); the pooled size-2-4 rate is 1.27%, partly reflecting deterministic even-size voting behavior. Second, a joint-correctness proxy (strict diversity) is nearly collinear with one minus mean accuracy (size-3 Spearman rho = +0.991 / +0.988); raw diversity-gain associations are strongly capability-entangled and, with one exception, unstable under control. Third, three linear contingency-table statistics are algebraically non-separable; after capability control, the empirically stable remainder is a modest residual pairwise co-failure association in which more shared error corresponds to lower gain. This direction is robust, but its magnitude is configuration-dependent. Joint rawspace linear regressions treating strict diversity, disagreement, and double-fault as independent predictors are rank-deficient by construction.