Search papers, labs, and topics across Lattice.
This study evaluates the dimensionality and measurement precision of the Humanity's Last Exam (HLE) by analyzing 29 large language models (LLMs) on its multiple-choice subset. Using a two-parameter logistic item response theory (IRT) model, the authors find that HLE primarily measures a single general reasoning factor, with domain labels contributing minimally to item response variance and showing high redundancy with total scores. Additionally, the analysis reveals that measurement precision is concentrated at moderate ability levels, indicating that HLE struggles to effectively differentiate between top-performing models.
HLE's domain-specific scores are largely redundant, revealing that its labels fail to capture distinct capabilities among leading language models.
Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text-only multiple-choice subset of HLE ($J = 428$ items) and apply psychometric methods to assess both the dimensionality of the benchmark and the distribution of its measurement precision. Fitting a two-parameter logistic IRT model, we find convergent evidence that HLE measures a single general reasoning factor: McDonald's $\omega_h = 0.998$, domain labels explain only 3.5\% of item response variance, within- and between-domain residual correlations are nearly identical (Cohen's $d = 0.016$), and domain-specific ability estimates are near-redundant with the total score ($r \geq 0.81$). A separate analysis of the test information function reveals that measurement precision concentrates at moderate ability levels and drops sharply above $\theta = 0$, where frontier models sit. These findings suggest that HLE's domain subscores do not warrant distinct capability interpretations and that the benchmark's ability to discriminate among the strongest models is limited.