Search papers, labs, and topics across Lattice.
This study investigates the limitations of traditional single-pass annotation methods for detecting factual errors in chatbot-generated medical responses, revealing that first-pass annotators often overlook subtle inaccuracies. By employing a multi-perspective approach that includes LLM-assisted candidate discovery and expert adjudication, the authors demonstrate that while LLMs can aid in identifying errors, they are not sufficient on their own. The findings indicate that multi-pass adjudication enhances the completeness of error detection benchmarks, but the quality of these benchmarks remains contingent on the expertise and judgment applied during the evaluation process.
Single-pass annotation can miss critical factual errors in chatbot responses, but a multi-perspective approach reveals a more comprehensive picture of accuracy.
Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect these errors is critical for determining chatbot safety. Yet factual-error detection is often treated as a single-pass, single-annotator labeling problem. In long-form chatbot responses, factual errors can be subtle and embedded within mostly correct text. We develop a multi-perspective annotation study of medically relevant chatbot responses, combining first-pass annotation, LLM-as-a-Judge (LaJ) candidate discovery, and two forms of adjudication: medical-expert and evidence-based fact-checking. First-pass annotators frequently miss factual errors later validated by adjudicators. LaJ improves candidate discovery, but is insufficient on its own: It misses factual errors that annotators catch. We also find disagreement among adjudicators, suggesting that adjudication over multiple candidate sources can improve benchmark completeness, but does not eliminate the need to apply judgment and expertise. Applied to an existing benchmark, this technique reveals a similar pattern of missing annotations. Together, these results suggest that in the settings examined here, single-pass hallucination benchmarks may achieve scale at the cost of undercounting factual errors. Multi-pass adjudication can improve coverage, but inferences drawn from the benchmarks are still sensitive to the judgment, expertise, and evidence used to determine error presence.