Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature
This paper compares specialized supervised Extreme Multi-Label Classification (XMLC) methods with lexical matching baselines and LLM-based methods for subject indexing contemporary German scientific literature.
The study provides a comprehensive evaluation of specialized supervised XMLC methods, lexical matching baselines, and LLM-based methods for subject indexing contemporary German scientific literature.
Keywords
Before reading this…
Applications
- →Library science
- →Information Retrieval
To understand this paper, make sure you know these concepts first:
- Understanding of Multi-Label Classificationfind papers →
- Knowledge of Extreme Multi-Label Classificationfind papers →
- Familiarity with library science conceptsfind papers →
Abstract
More Like ThisWith a large controlled vocabulary as the label set, the task of automated subject indexing in a library can be understood as a multi-label classification task. If the set of subject terms is large, the problem fits the Extreme Multi-Label Classification (XMLC) objective. In this study, we apply a selection of specialised supervised XMLC methods to the test case of subject indexing contemporary German scientific literature, collected at the German National Library (DNB). We contrast these results by including a classical lexical matching baseline and three of our own recently developed LLM-based methods into the benchmark. Algorithms are evaluated and compared in several metrics. This includes binary relevance comparisons with previously indexed material, as well as graded relevance ratings by professional subject librarians. A challenge for all methods is to reliably make suggestions from the long tail of the subject vocabulary. We find that supervised XMLC algorithms relying on transformer-based dense features give best results in terms of overall binary relevance metrics. However, focusing on graded relevance and performance in the long tail of our subject vocabulary, the LLM-based generative methods give better results, making them a promising alternative for future productive use.