Search papers, labs, and topics across Lattice.
VetScore is a novel multi-step evaluation method designed for assessing the reliability of generated claims in veterinary long-form question answering by incorporating citation excerpts and weighing claims by their potential harm. The method segments outputs into individual claims, scores them based on their faithfulness to source excerpts and their associated risk, and calculates an overall risk-adjusted score. Evaluation results demonstrate that VetScore aligns closely with expert veterinary assessments, even when using smaller judge models, highlighting its effectiveness and explainability.
VetScore reveals that risk-weighted evaluations can significantly enhance the reliability of veterinary QA systems, achieving expert-level alignment even with minimal model complexity.
Citation excerpts can be used to increase the reliability of generated outputs and their faithfulness to cited sources, which is especially important in high-stakes domains such as human and veterinary medicine. However, this does not guarantee that generated claims are faithful to the provided excerpts. We present VetScore, a multi-step evaluation method for veterinary long-form question answering, designed to assess how well are generated claims supported by the provided excerpts, weighing this information by each claim's harm potential. VetScore first segments the output and decomposes it into individual claims, then scores each claim with respect to its harm potential and evaluates its faithfulness to source excerpts, and finally calculates the overall risk-adjusted score. We collect an expert-annotated meta-evaluation dataset, evaluate our approach with a range of judge models, and show that it achieves high correlations with veterinary experts even with small judge models, while offering explainability across multiple dimensions.