Open Korean Corpora: A Practical Report
Autor: | Cho, Won Ik, Moon, Sangwhan, Song, Youngsook |
---|---|
Rok vydání: | 2020 |
Předmět: | |
Druh dokumentu: | Working Paper |
DOI: | 10.18653/v1/2020.nlposs-1.12 |
Popis: | Korean is often referred to as a low-resource language in the research community. While this claim is partially true, it is also because the availability of resources is inadequately advertised and curated. This work curates and reviews a list of Korean corpora, first describing institution-level resource development, then further iterate through a list of current open datasets for different types of tasks. We then propose a direction on how open-source dataset construction and releases should be done for less-resourced languages to promote research. Comment: Published in NLP-OSS @EMNLP2020; May 2023 version added with new datasets |
Databáze: | arXiv |
Externí odkaz: |