Search papers, labs, and topics across Lattice.
This paper leverages representation theorems from decision theory to establish a label-free framework for evaluating and regularizing large language models (LLMs) and AI systems based on their compliance with rationality axioms. By checking a model's responses to synthetic choice problems, the authors demonstrate that rationality can be assessed without external labels, providing computable penalties for non-compliance. The findings suggest that models passing these checks cannot be rejected on rationality grounds by any further tests, thus offering a robust method for ensuring rational behavior in AI systems.
Label-free evaluation of AI systems reveals that models can be rigorously assessed for rationality without relying on external labels or human feedback.
Representation theorems in decision theory establish that behavior satisfies certain axioms if and only if it can be rationalized by a well-defined objective. I argue that this ``if and only if'' structure provides a potentially useful foundation for label-free evaluation and regularization of LLMs and other AI systems. Axiom compliance can be checked from the model's own responses to synthetic choice problems, with no external labels or human feedback, and the penalties are readily computable. Because the axioms are necessary and sufficient, the resulting checks exhaust the implications of the relevant rationality standard for the elicited data: a model that passes cannot be rejected on rationality grounds by any further test of the same data. I discuss three instantiations: probabilistic coherence via a theorem of de Finetti, preference rationality via Afriat's theorem, and subjective expected utility via a theorem of Echenique and Saito (2015), each yielding a continuous penalty that is zero whenever behavior can be rationalized. Since coherence does not restrict which objective rationalizes behavior, these penalties complement rather than replace other evaluation and training signals.