Annif Analyzer Shootout: Comparing text lemmatization methods for automated subject indexing

Autor:	Osma Suominen, Ilkka Koskenniemi
Jazyk:	angličtina
Rok vydání:	2022
Předmět:	Bibliography. Library science. Information resources
Zdroj:	Code4Lib Journal, Iss 54 (2022)
Druh dokumentu:	article
ISSN:	1940-5758
Popis:	Automated text classification is an important function for many AI systems relevant to libraries, including automated subject indexing and classification. When implemented using the traditional natural language processing (NLP) paradigm, one key part of the process is the normalization of words using stemming or lemmatization, which reduces the amount of linguistic variation and often improves the quality of classification. In this paper, we compare the output of seven different text lemmatization algorithms as well as two baseline methods. We measure how the choice of method affects the quality of text classification using example corpora in three languages. The experiments have been performed using the open source Annif toolkit for automated subject indexing and classification, but should generalize also to other NLP toolkits and similar text classification tasks. The results show that lemmatization methods in most cases outperform baseline methods in text classification particularly for Finnish and Swedish text, but not English, where baseline methods are most effective. The differences between lemmatization methods are quite small. The systematic comparison will help optimize text classification pipelines and inform the further development of the Annif toolkit to incorporate a wider choice of normalization methods.
Databáze:	Directory of Open Access Journals
Externí odkaz:	https://doaj.org/article/5cb9e0e3a3e4476b8b6ca14d15efdb72 Zobrazit plný text záznamu View record in DOAJ Plný text ve formátu HTML
Nepřihlášeným uživatelům se plný text nezobrazuje	K zobrazení výsledku je třeba se přihlásit.