Search papers, labs, and topics across Lattice.
This paper presents an open-source pipeline, Institutional Books - Visual Elements, designed to extract, classify, deduplicate, and caption visual elements from historical book collections, addressing the underutilization of these components in digitization projects. By leveraging this pipeline, the authors provide a dataset of 22.6 million visual elements sourced from nearly a million scanned volumes, significantly enhancing the accessibility of rich visual content for various research applications. The findings underscore the potential for automated workflows to unlock new avenues in digital humanities and AI model training, which have been historically limited to textual data.
Unlocking 22.6 million visual elements from historical books could revolutionize how we engage with digitized library collections and their applications in AI and research.
Historical book collections contain rich visual elements - such as illustrations, photographs, engravings, and decorative art - that are frequently under-explored in large-scale digitization projects. While Optical Character Recognition (OCR) has standardized the extraction of textual content, these visual components offer a layer of nuance and context that remains largely untapped by automated text extraction workflows. This technical report introduces Institutional Books - Visual Elements, an open-source end-to-end pipeline for detecting, classifying, deduplicating, and captioning visual elements from historical book collections. Alongside this pipeline, we release an initial dataset of 22.6 million visual elements extracted from the 983,004 scanned volumes that comprise the Institutional Books: Harvard Library dataset. This work contributes to ongoing, community-wide efforts to enable new use cases for digitized library collections through computational access, from artificial intelligence model training to digital humanities research.