Search papers, labs, and topics across Lattice.
The authors extend the statistical physics Potts model to evaluate multi-category scoring reliability by parameterizing pairwise rater-category agreement indicators with LLM-derived embedding similarities. Unlike standard psychometric and evaluation frameworks that enforce rigid equidistant scales or threshold assumptions, this formulation models multi-rater concordance directly through semantic feature spaces. Across balanced short-answer and imbalanced essay benchmarks, the approach aligns closely with human raters, restricting errors almost exclusively to adjacent score bands via a tunable similarity-power transformation.
Statistical physics meets LLM evaluation: mapping multi-category rater agreement onto a Potts model preserves true rubric ordinality without requiring rigid threshold or distance assumptions.
The Ising model is extended to the Potts model for multinomial data. We introduce a Rater Ising-Potts model that uses agreement indicators between pairs of raters and category labels, with weights derived from LLM embeddings. The model does not presuppose ordered category thresholds or equidistant scoring; instead, it focuses directly on pairwise agreement among raters and assigns category-specific positive weights, making it particularly suited for multi-category scoring reliability when raters evaluate responses using a scoring guide. We demonstrate the model's effectiveness on diverse constructed-response tasks, including balanced short-answer items and more challenging, imbalanced essay prompts from the AERA dataset. Across these settings, the model achieves strong agreement with human scores, with the vast majority of misclassifications occurring between adjacent score levels, confirming its ability to preserve the ordinal structure of scoring rubrics without imposing rigid assumptions. A practical similarity normalization and optional power transformation is introduced as a tunable preprocessing step that sharpens semantic distinctions and can be adapted to different datasets. These findings suggest that LLM-derived semantic similarities, combined with this parsimonious Potts-type formulation and flexible similarity scaling, offer a robust and interpretable framework for reliability auditing in educational assessment contexts. Extensions to multiple raters and hierarchical rating processes are discussed.