A Comparison of Pre-Trained Language Models for Multi-Class Text Classification in the Financial Domain

Autor:	Jacques Klein, Anne Goujon, Kevin Allix, Tegawendé F. Bissyandé, Lisa Veiber, Cedric Lothritz, Yusuf Arslan
Jazyk:	angličtina
Rok vydání:	2021
Předmět:	Finance Computer science [C05] [Engineering computing & technology] Vocabulary Class (computer programming) Leverage (finance) Artificial neural network Computer science business.industry media_common.quotation_subject Sciences informatiques [C05] [Ingénierie informatique & technologie] Financial Text Classification Domain (software engineering) Task (project management) Benchmark (surveying) Language model business media_common
Zdroj:	WWW (Companion Volume)
Popis:	Neural networks for language modeling have been proven effective on several sub-tasks of natural language processing. Training deep language models, however, is time-consuming and computationally intensive. Pre-trained language models such as BERT are thus appealing since (1) they yielded state-of-the-art performance, and (2) they offload practitioners from the burden of preparing the adequate resources (time, hardware, and data) to train models. Nevertheless, because pre-trained models are generic, they may underperform on specific domains. In this study, we investigate the case of multi-class text classification, a task that is relatively less studied in the literature evaluating pre-trained language models. Our work is further placed under the industrial settings of the financial domain. We thus leverage generic benchmark datasets from the literature and two proprietary datasets from our partners in the financial technological industry. After highlighting a challenge for generic pre-trained models (BERT, DistilBERT, RoBERTa, XLNet, XLM) to classify a portion of the financial document dataset, we investigate the intuition that a specialized pre-trained model for financial documents, such as FinBERT, should be leveraged. Nevertheless, our experiments show that the FinBERT model, even with an adapted vocabulary, does not lead to improvements compared to the generic BERT models.
Databáze:	OpenAIRE
Externí odkaz:	https://explore.openaire.eu/search/publication?articleId=doi_dedup___::f9cc57ba24de8d8eb7eaca5c95bcd0cd http://orbilu.uni.lu/handle/10993/47288 Zobrazit plný text záznamu