Search papers, labs, and topics across Lattice.
This study benchmarks various methods for automated subject indexing of German scientific literature, specifically focusing on Extreme Multi-Label Classification (XMLC) with a large controlled vocabulary. By comparing specialized supervised XMLC methods, classical lexical matching, and novel LLM-based approaches, the research reveals that while transformer-based supervised algorithms excel in binary relevance, LLM-based methods outperform in graded relevance and handling the long tail of subject terms. These findings suggest that generative AI can effectively enhance the indexing process, offering a viable alternative to traditional supervised methods.
LLM-based methods outperform traditional supervised algorithms in indexing nuanced subject terms, particularly in the long tail of vocabulary.
With a large controlled vocabulary as the label set, the task of automated subject indexing in a library can be understood as a multi-label classification task. If the set of subject terms is large, the problem fits the Extreme Multi-Label Classification (XMLC) objective. In this study, we apply a selection of specialised supervised XMLC methods to the test case of subject indexing contemporary German scientific literature, collected at the German National Library (DNB). We contrast these results by including a classical lexical matching baseline and three of our own recently developed LLM-based methods into the benchmark. Algorithms are evaluated and compared in several metrics. This includes binary relevance comparisons with previously indexed material, as well as graded relevance ratings by professional subject librarians. A challenge for all methods is to reliably make suggestions from the long tail of the subject vocabulary. We find that supervised XMLC algorithms relying on transformer-based dense features give best results in terms of overall binary relevance metrics. However, focusing on graded relevance and performance in the long tail of our subject vocabulary, the LLM-based generative methods give better results, making them a promising alternative for future productive use.