Search papers, labs, and topics across Lattice.
To eliminate the fragile assumption that reference "anchors" in LLM-as-a-judge panels are error-free, this work derives a closed-form estimator and diagnostic battery to point-identify latent quality variance, common-mode judge variance, and anchor contamination correlation. This addresses a critical failure mode where standard clean-anchor estimators report contaminated companion anchors as pristine whenever the trusted anchor is compromised, alongside proving that identification collapses under purely ordinal scores without at least three continuous anchors. While synthetic experiments validate parameter recovery, model-adequacy pre-tests consistently and correctly reject real-world judge panels, establishing concrete mathematical boundaries for when common-factor evaluation assumptions fail.
Standard LLM evaluation pipelines blindly trust reference "anchors", but a single contaminated benchmark causes conventional estimators to falsely certify an entire panel's corrupted companion anchors as clean.
When an external reference set (an anchor) is used to decompose an LLM-judge panel's error into a quality signal and a shared common-mode error, standard practice assumes the anchor is uncontaminated: its error uncorrelated with the judges' shared error. We study when that assumption can be dropped and replaced by an estimate. Under a single-common-factor model, >=2 judges and >=2 anchors point-identify the quality variance, the common-mode variance, and each anchor's contamination correlation rho_k in closed form, with an exact per-anchor-pair failure boundary; a designated clean-anchor estimator, by contrast, reports a contaminated companion anchor as fully clean once its trusted anchor is itself contaminated. Because the single-common-factor assumption is itself untestable, the estimator ships gated behind a calibrated diagnostic battery (judge-covariance dispersion; over-identification; a family-block test from judge metadata, with a family-blocked estimator that removes family-level shared-residual bias exactly), bootstrap confidence intervals with measured coverage, and a weak-identification screen. A proposition maps which violations bias rho_k, in which direction, and which evade detection. For ordinal scores we show an identification hierarchy: with all variables ordinal, rho_k is not identified at any number of anchors; with ordinal judges and >=3 continuous anchors it is, and we give an estimator for that case. On real data the validation is asymmetric, and we say so plainly: the diagnostics are validated in the rejecting direction (both real panels we test are correctly rejected by the model-adequacy pre-test), while the estimator is validated in simulation and stress-tested semi-synthetically under oracle calibration; no real panel has yet passed the pre-test, and the pre-test exists precisely to say so. All results replay offline from shipped, checksummed artifacts.