Search papers, labs, and topics across Lattice.
This paper presents a machine-readable catalogue for the personal archive of Konstantin Tsiolkovsky, comprising 2,019 files and 51,008 scans, along with a method for assessing the accuracy of handwritten text recognition in the absence of ground truth. By leveraging the dual preservation of texts as manuscripts and typed copies, the authors establish a framework to quantify reading errors, revealing that median agreement between two readings of handwritten pages is only 37%. Furthermore, the catalogue's accuracy is validated against published editions, demonstrating a high correlation and providing insights into the limitations of collating redacted texts.
Handwritten text recognition in historical archives reveals a surprising 63% disagreement even between expert transcriptions of the same page.
The personal archive of Konstantin Tsiolkovsky (1857-1935) is held as fond 555 of the Archive of the Russian Academy of Sciences. The archive scanned the fond and published the images, but with no queryable catalogue, no full-text search and no dataset: the holdings can only be browsed one page at a time. This paper describes a machine-readable catalogue of all 2,019 files and 51,008 scans, a dating for 1,969 files taken from the archive's own descriptions, a page-level classification of every scan into handwriting and typescript, and a growing corpus of machine transcriptions (currently 322 files, 5,454 scans). It also reports a way to measure handwritten-text-recognition accuracy in an archive with no ground truth. Archives of the typewriter era often preserve one text twice, as manuscript and as a typed copy; transcribing both and comparing isolates the reading error, since source and pipeline are identical and only page difficulty differs. Across 294 such pairs from 27 files, two readings of a handwritten page agree on a median 37% of words. On two files that also have a published edition the estimate can be checked against ground truth: it is unbiased to within a percentage point and ranks pages as the truth does (rank correlation 0.92 where the edition is a faithful witness). This bounds use: two variants of one work here share 19% of words, below the rate at which two readings of a single page agree, so the redactions cannot be collated word by word at this quality. That negative result is reported as such, and the constraint is built into the tool.