Search papers, labs, and topics across Lattice.
This paper introduces LegalPincite, a comprehensive legal information retrieval dataset designed to enhance the evaluation of legal IR methods by addressing the shortcomings of existing datasets. It features masked case and paragraph queries, a complete corpus of all paragraphs, and validated ground-truth citations, enabling a more realistic retrieval setting without data leakage. The dataset supports multi-level retrieval tasks, thereby providing a robust framework for developing and assessing legal IR systems.
LegalPincite reveals that existing legal IR datasets can mislead performance evaluations due to data leakage, offering a more accurate foundation for legal information retrieval research.
A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While legal practice often requires pinpoint citations (pincites) to specific case paragraphs, most existing public legal IR datasets lack paragraph-level citation annotations. Yet, publicly available datasets with such information contain data leakage in the query text and exclude paragraphs that are neither citing nor cited from the corpora, creating an unrealistic and oversimplified retrieval setting, potentially leading to inflated performance. To address these limitations, we contribute a large-scale legal IR dataset constructed from Court of Justice of the European Union (CJEU) judgments. The dataset contains: (i) masked case/paragraph queries, with removed citation information; (ii) a corpus that includes all paragraphs; and (iii) case- and paragraph-level ground-truth citations, with partial human expert validation. Our dataset supports both the development and rigorous evaluation of legal IR methods, at multiple query-document levels (case-to-case, paragraph-to-case, and paragraph-to-paragraph retrieval). Link to dataset: https://huggingface.co/datasets/theresiavr/legalpincite