Search papers, labs, and topics across Lattice.
This study investigates the reliance of large language models (LLMs) on proxy attributes in decision-making tasks, specifically in clinical rankings. By measuring causal proxy effects across four LLMs, the authors reveal that models often exhibit uncalibrated reliance on proxies, leading to potential discrimination or sound inference depending on the context. The findings indicate that reliance on informative proxies can be misaligned with the evidence, with one model showing a significant drop in reliance when social field names are used, highlighting the fragility of social-label suppression.
LLMs can misjudge the relevance of proxy attributes, leading to uncalibrated decision-making that risks both discrimination and erroneous inference.
Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carry predictive value, so a changed decision can be discrimination or sound inference. We measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth, where the reliance the evidence warrants can be computed exactly and used as the reference. One audit signal yields three verdicts: over-reliance, warranted and under-reliance. Under neutral labels every model relies on proxies with no information. Informative proxies draw all three. Social field names push reliance down, below the reference in one model. Two findings explain this. Reliance severely undertracks the evidence, and social-label suppression is fragile, since in-context examples raise it above zero in every model. Accuracy-based evaluation detects none of this.