Search papers, labs, and topics across Lattice.
This paper critiques the current methods for evaluating epistemic uncertainty, highlighting a misalignment between Bayes-optimal decision strategies and commonly used uncertainty scores. By leveraging the epistemic reject-option framework, the authors formulate selective prediction as a constrained optimization problem, proving that the optimal selector is a thresholded convex combination of aleatoric and epistemic uncertainties. Their findings reveal significant discrepancies between decision-theoretic rankings and traditional proxy-task rankings, suggesting that standard correlation metrics may not accurately reflect operational utility in uncertainty disentanglement.
Decision-theoretic evaluations of epistemic uncertainty reveal that traditional metrics can mislead researchers about the effectiveness of their models.
Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to quantify epistemic uncertainty. Building on the epistemic reject-option framework, we evaluate epistemic uncertainty using its ability to identify regret, the reducible error. Formulating selective prediction as a constrained optimization over coverage, expected risk, and regret, we prove the optimal selector is a thresholded convex combination of the ground-truth aleatoric and epistemic uncertainties. This theoretical unification exposes a weakness in recent uncertainty disentanglement literature: we demonstrate that standard correlation metrics between learned components do not necessarily predict their actual operational utility. We instead propose to evaluate the achievable risk, regret, coverage surface of the decomposition as a diagnostic for joint disentanglement and utility. Benchmarking standard methods on datasets with dense human annotations reveals that decision-theoretic rankings can disagree substantially with proxy-task rankings, including pairwise rank inversions between methods that are top-ranked on one criterion and bottom-ranked on other.