Search papers, labs, and topics across Lattice.
This study investigates the challenges of auditing multiple LLM agents with a limited budget of audits, revealing that miscalibrated self-reported confidence can lead to suboptimal auditing outcomes. By modeling the auditing process using a two-level Gaussian copula, the authors identify a critical miscalibration threshold, $\delta^*$, which increases as the audit budget decreases, indicating that confidence-ranked auditing can perform worse than random selection under certain conditions. The findings show that while some models exhibit near-constant confidence levels, a proprietary model demonstrates more informative confidence, highlighting the importance of model selection in effective oversight.
Confidence-ranked auditing of LLM agents can backfire, with miscalibration thresholds rising as audit budgets shrink, leading to worse outcomes than random selection.
A single human must audit $N$ LLM agents under a budget of $B \ll N$ audits per round, guided by self-reported confidence that may be adversarially miscalibrated and by correlated errors. We model this as budgeted noisy inspection over a two-level Gaussian copula and locate the miscalibration threshold $\delta^*$ past which confidence-ranked auditing is \emph{worse} than random. Two a-priori expectations reverse: $\delta^*$ \emph{rises} as the budget shrinks, and cross-family correlation is not low---shared difficulty dominates lineage. Five open-weight LLMs show operationally useless (near-constant) confidence, point estimates at or beyond the flip though CIs straddle it; a proprietary model is informative and lands below it. We give a quantitative criterion for \emph{vacuous} oversight, and replaying policies on recorded traces confirms the ordering.