Search papers, labs, and topics across Lattice.
HIPE-2026 advances the extraction of person-place relations from noisy multilingual historical texts, focusing on two temporally grounded relation types: $at$ and $isAt$. The evaluation campaign engaged 17 teams to tackle challenges such as historical language variation and OCR noise, utilizing datasets from 19th and 20th-century newspapers and early modern French literature. Results from over 40 submissions demonstrate a spectrum of strategies, revealing critical trade-offs between accuracy, computational efficiency, and robustness in historical relation extraction.
Extracting historical person-place relations reveals that lightweight classifiers can outperform large language models in specific contexts, challenging assumptions about model size and performance.
Was this person ever at that place, and if so, when? Answering such questions from noisy, multilingual historical documents is the central challenge of HIPE-2026, the third edition of the HIPE evaluation series. Moving from named entity recognition and linking (HIPE-2020, HIPE-2022) to reasoning about relationships between entities, HIPE-2026 targets two temporally grounded relation types: $at$, indicating that a person was present at a location at some point prior to a document's publication date, and $isAt$, indicating presence contemporaneous with that date. This paper presents the results of the evaluation campaign, which confronted 17 participating teams with the challenges of historical language variation, OCR noise, and indirect contextual cues across three languages: French, German, and English. The datasets include historical newspaper text from the nineteenth and twentieth centuries, as well as a surprise-domain generalization set drawn from early modern French literary texts. A distinctive feature of HIPE-2026 is its three-fold evaluation framework, which assesses predictive accuracy, computational efficiency, and cross-domain generalization, reflecting the practical demands of large-scale historical document processing in the cultural heritage domain. Across more than 40 submitted runs, results reveal a wide range of strategies, from state-of-the-art large language models to lightweight task-specific classifiers, and highlight the trade-offs between accuracy, efficiency, and robustness inherent to historical relation extraction at corpus scale. System descriptions, datasets, and findings are presented and discussed, offering a detailed picture of the current state of temporally grounded relation extraction for historical documents.