Search papers, labs, and topics across Lattice.
This study investigates soft inter-rater reliability in unstructured biomedical text annotations, addressing the challenge of quantifying agreement among annotators using open-ended labels. Through synthetic experiments, the authors demonstrate that various semantic equivalence measures can effectively quantify this reliability, revealing that the choice of measure significantly influences the estimation's failure modes. The findings indicate that while embeddings offer scalability, they struggle with nuanced distinctions, and large language models face scalability issues for chance agreement estimation, leading to the recommendation of natural language inference-based measures as a viable solution.
Semantic equivalence measures can drastically change how we quantify annotator agreement in biomedical text, revealing hidden biases in existing methods.
Biomedical literature is increasingly mined for knowledge beyond the questions it was written to answer. Because the target concepts are not known in advance, annotators prefer open-ended labels, whose agreement is hard to quantify. We study soft inter-rater reliability for annotators providing unstructured texts for biomedical annotation tasks. Synthetic experiments show that soft reliability can be quantified using a variety of semantic equivalence measures, and that the choice of measure affects failure modes of the estimation. Embeddings are scalable, but limited when differentiating similar but distinct concepts. Large language models are promising, but limited by scalability for estimating agreement by chance. Finally, we suggest measures based on natural language inference as a sensible compromise.