Extracting Data from Comparable Corpora

Autor:	Mārcis Pinnis, Nikola Ljubešić, Inguna Skadiņa, Tatjana Gornostaja, Špela Vintar, Darja Fišer, Marko Tadić, Dan Stefanescu
Rok vydání:	2019
Předmět:	Data extraction Machine translation business.industry Computer science Order (business) Search engine indexing Question answering Artificial intelligence business computer.software_genre computer Reciprocal Natural language processing
Zdroj:	Using Comparable Corpora for Under-Resourced Areas of Machine Translation ISBN: 9783319990033 Using Comparable Corpora for Under-Resourced Areas of Machine Translation
Popis:	Comparable corpora may comprise different types of single-word and multi-word phrases that can be considered as reciprocal translations, which may be beneficial for many different natural language processing tasks. This chapter describes methods and tools developed within the ACCURAT project that allow utilising comparable corpora in order to (1) identify terms, named entities (NEs), and other lexical units in comparable corpora, and (2) to cross-lingually map the identified single-word and multi-word phrases in order to create automatically extracted bilingual dictionaries that can be further utilised in machine translation, question answering, indexing, and other areas where bilingual dictionaries can be useful.
Databáze:	OpenAIRE
Externí odkaz:	https://explore.openaire.eu/search/publication?articleId=doi_________::c8f31fc81d3de2dcbe75b06e778fbe37 https://doi.org/10.1007/978-3-319-99004-0_4 Zobrazit plný text záznamu