Search papers, labs, and topics across Lattice.
This paper introduces a novel framework for evaluating Large Language Models (LLMs) by employing controlled factorial experiments to isolate causal effects and utilizing exact token-level Probability Mass Functions (PMFs) to eliminate sampling noise. By integrating principles from human psychometrics, the authors derive a multivariate ordinal consensus metric and distributional ANOVA to analyze biases in LLMs more precisely. A case study on consumer ethnocentrism across five LLMs illustrates the framework's effectiveness in revealing country-of-origin biases that traditional benchmarks fail to distinguish.
Systematic evaluation reveals that traditional benchmarks obscure significant biases in LLMs, which can be isolated through a new analytical framework.
As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks. While effective for general capabilities, these datasets fundamentally conflate causal mechanisms: even when an aggregate bias is detected, unstructured evaluations cannot disentangle whether it stems from baseline traits, contextual confounders, or complex interactions. To address this, we introduce an analytically exact framework for the controlled behavioral evaluation of LLMs. We bridge human psychometrics with LLM mechanics by resolving gaps in design, measurement, and analysis. First, we replace unstructured prompting with fully crossed factorial experiments to systematically isolate causal main and interaction effects. Second, we eliminate Monte Carlo text sampling noise by operating directly on exact, token-level Probability Mass Functions (PMFs). Third, we derive a multivariate ordinal consensus metric and a distributional ANOVA to process these PMFs analytically. We validate our framework with a case study on consumer ethnocentrism across five LLMs, demonstrating how our approach isolates systemic country-of-origin biases that aggregate benchmarks otherwise obscure.