Search papers, labs, and topics across Lattice.
To determine whether LLM evaluations track their intended constructs, this work applies psychometric convergent and discriminant validity alongside item response theory across 56 capability and safety benchmarks evaluated on 53 models. The analysis reveals that safety benchmarks lack internal consistency, capability benchmarks fail to discriminate between distinct concepts like reasoning and knowledge, and shared design artifacts such as scoring formats often dictate correlations more than underlying targets. These findings highlight that many widely used metrics do not measure what they claim, evidenced by bias benchmarks like BBQ-accuracy correlating more strongly with reasoning suites than with other bias evaluations.
Standard AI benchmarks often measure evaluation formatting rather than their purported constructs鈥攚ith "reasoning" and "knowledge" metrics failing to discriminate from one another, and bias benchmarks correlating more strongly with general reasoning than with peer bias tests.
Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We adapt convergent and discriminant validity from the social sciences into an approach for interrogating AI benchmarks, applying it to 56 capability and safety benchmarks across 53 models. We label benchmarks with substantively similar purported concepts to a shared assigned concept, and ask whether model rankings on benchmarks with the same assigned concept correlate more strongly than rankings on benchmarks with different assigned concepts. We ask analogous questions at the item level using item response theory (IRT) models. We find that correlations between model rankings on benchmarks with the same assigned safety concepts are often weak, suggesting these concepts may be conceptualized inconsistently across benchmarks. For assigned capability concepts (e.g., reasoning, knowledge), model rankings are often as strongly correlated among benchmarks with the same assigned concept as between benchmarks with different assigned concepts, suggesting these capability concepts may not discriminate well from one another. In some cases, benchmarks that share design elements (e.g., score format) correlate more strongly than benchmarks with the same assigned concept. Finally, some individual benchmarks correlate more strongly with benchmarks assigned a different concept than with benchmarks sharing their own assigned concept, suggesting they may measure a different concept than they purport to. For example, BBQ-accuracy correlates more strongly with benchmarks labeled reasoning than with benchmarks that share its assigned concept, bias. To support future empirical work on benchmark validity, we release our extensive dataset of model outputs and scores at the item- and benchmark-level.