Search papers, labs, and topics across Lattice.
This paper critiques the use of difference-in-differences designs in audits of LLM judges by revealing that the bounded rating scale can obscure true effects due to differential attenuation. The authors demonstrate that the observed interactions in their pre-registered audit of a pedagogy judge stem from a severity shift rather than genuine preference differences, leading to a misleading interpretation of the results. Their findings indicate that a nominally significant interaction can be largely attributed to the scale's inherent limitations, rather than actual bias in the judge's ratings.
A seemingly significant interaction in LLM judge audits may be a mirage, driven by scale limitations rather than true preference differences.
Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them. We exhibit the failure inside a pre-registered audit of a frozen pedagogy judge, sealed before the first of its 990 calls. The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\% BCa $[-0.167, +0.353]$, $p = 0.684$). The audit's one nominally significant interaction, $+0.378$ ($p = 0.002$), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85\% of it from the observed severity shift and the scale floor alone. We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings.