Search papers, labs, and topics across Lattice.
This paper introduces dilemmadata, a novel resource that reconciles the AugmentedNet Dataset and the Distant Listening Corpus to create a unified, interoperable corpus of Roman numeral analyses. The authors address four key dilemmata鈥攁nnotation-standard, representational, toolchain, and curatorial鈥攂y transforming and augmenting data while preserving musical semantics, resulting in a comprehensive dataset of 1,621 pieces and approximately 2.8 million note-wise annotations. The dataset not only facilitates comparative analysis between different analytical traditions but also serves as a foundation for refining Roman-numeral encoding standards in future research.
The largest homogeneous Roman-numeral corpus ever created allows for direct note-for-note comparison across distinct analytical traditions, challenging existing standards in tonal harmony research.
In recent years, there has been growing effort to annotate and collect large-scale corpora of Roman numeral analyses in support of data-driven studies in tonal harmony. We introduce dilemmadata, the first resource to reconcile two major collections, the AugmentedNet Dataset (AN) and the Distant Listening Corpus (DLC), making them interoperable through a shared note-wise TSV schema. The reconciliation confronts four families of dilemmata: annotation-standard (the two encode the same musical fact differently in terms of vocabulary size, syntax, conventions for chord extensions, inventory of special chord functions), representational (what counts as a row, and which information survives the conversion), toolchain (incompatible Python ecosystems built around music21 vs. ms3+dimcat), and curatorial (which pieces to include, exclude, or retain twice). We resolve each by deliberately transforming, augmenting, and omitting information, formalising the mismatches, preserving musical semantics, and flagging transformations that may subtly affect annotation fidelity. Consistency checks and qualitative inspections offer a preliminary assessment of post-conversion validity and a basis for critiquing the theoretical assumptions embedded in each original standard. After removing duplicates and merging the two collections, the resulting dilemmadata (1,621 pieces and aprox. 2.8 M note-wise annotations) is the largest homogeneous Roman-numeral corpus currently available, albeit far from perfect. Crucially, we retain 84 pieces common to both corpora under each of their original analyses, yielding a shared reference set in which two equally legitimate analytical traditions can be compared note-for-note over identical musical material. Released on Zenodo, dilemmadata supports interoperability, comparative harmonization modeling, and future refinement of Roman-numeral encoding standards.