Search papers, labs, and topics across Lattice.
This paper introduces a self-evolving critic framework called Critic Experience Bank (CEB) that enables LLM agents to estimate step-level confidence by accumulating evidence from past actions and their outcomes. By utilizing a hindsight LLM to evaluate the productivity of actions based on full execution feedback, CEB dynamically updates a memory bank of experiences that informs future decision-making. The method significantly improves calibration and ranking metrics across various benchmarks, achieving up to a 54% reduction in expected calibration error compared to the best existing training-free baseline.
A self-evolving critic can reduce confidence estimation errors in LLM agents by up to 54% without requiring any training or ground truth labels.
LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong step can waste interaction budget or trigger irreversible side effects long before the final failure is observed. Reliable deployment therefore requires \emph{step-level confidence estimation}: a calibrated probability that each proposed action is productive, available \emph{before} the action is executed. Existing LLM confidence estimators are designed to score a response from the given prompt, but agent confidence also depends on execution consequences: whether similar actions in similar situations actually advanced the task after the environment responded. We introduce the \method (\methodshort), a self-evolving critic framework in which an LLM critic accumulates evidence from its own past judgments and their observed consequences. After each trajectory, a hindsight LLM that sees the full execution feedback votes on whether each step was productive. The resulting pseudo-labels populate a memory bank from which related productive and unproductive experiences are retrieved into the critic's prompt whenever a similar step recurs. \methodshort requires no training and uses no ground truth step labels. Across three agent benchmarks and three critic backbones, \methodshort attains the best calibration (ECE and Brier) and ranking (AUC) in every dataset--critic combination, reducing ECE by up to $54\%$ relative to the strongest training-free baseline.