Search papers, labs, and topics across Lattice.
Across 40,726 pre-registered queries to five LLMs, the authors tested whether a published bias reversal鈥攚here models favor minority applicants in isolated ratings but penalize them in side-by-side rankings鈥攇eneralizes from charitable aid to hiring, lending, and medical triage. None of the 36 planned contrasts survived multiple testing corrections, conclusively bounding out the ranking penalty in hiring and cutting rating disparities in half compared to prior claims. Crucially, the evaluations revealed that instrument effects dominate the signal: models exhibited first-candidate position biases that rivaled or exceeded demographic disparities, and reliably tied identical content whether the varied attribute was race or a hobby.
Position bias and audit design rival or exceed demographic disparities in LLMs, rendering high-profile findings of rating-vs-ranking bias reversals non-replicable across hiring, lending, and triage.
Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lending, and medical triage: 40,726 requests to five models, applications differing only in the applicant's name, and a primary test fixed before collection. It does not. None of 36 planned contrasts survives correction. The rating advantage keeps its sign at roughly half the published size, and a precision extension bounds any hiring ranking penalty below the published effect, though the lending and triage ranking floors sit above that margin, so the exclusion is conclusive for hiring ranking and for rating in all three domains only. Planted disparities tracking their injected sizes and a directional replication on the original aid materials bound these nulls. The audit is livelier than the demographics: models recognize transparent audits nearly always, tie every identical-content comparison whether the varying detail is race or a hobby, and reward first-listed candidates as much as any demographic effect we measure. Audit verdicts reflect audit construction more than demographic bias.