Toward enhanced Arabic speech recognition using part of speech tagging

Autor:	Husni Al-Muhtaseb, Moustafa Elshafei, Dia AbuZeina, Wasfi G. Al-Khatib
Rok vydání:	2011
Předmět:	Linguistics and Language Speech production Stop words Computer science business.industry Speech recognition Word error rate Speech corpus computer.software_genre Language and Linguistics language.human_language Speech shadowing Human-Computer Interaction Compound Modern Standard Arabic language Computer Vision and Pattern Recognition Language model Artificial intelligence business computer Software Natural language processing
Zdroj:	International Journal of Speech Technology. 14:419-426
ISSN:	1572-8110 1381-2416
DOI:	10.1007/s10772-011-9121-5
Popis:	One major source of suboptimal performance in automatic continuous speech recognition systems is misrecognition of small words. In general, errors resulting from small words are much more than errors resulting from long words. Therefore, compounding some words (small or long) to produce longer words is welcome by speech recognition decoders. In this paper, we present a novel approach to artificially generate compound words using part of speech tagging. For this purpose, we consider two cases in Arabic speech where two words are pronounced without a silence period in between: a noun followed by an adjective, and a preposition followed by any word. To collect the candidate compound words, we use Stanford Arabic tagger to tag all words in our baseline transcription corpus. Then, compound words are generated whenever any of the two cases occur in a sequence of two words. The unique compound words are then added to the expanded pronunciation dictionary, as well as to the language model. Using Sphinx 3, we test the proposed method for a 5.4 hours speech corpus of modern standard Arabic. The results show a significant improvement, as the word error rate is reduced by 2.39%.
Databáze:	OpenAIRE
Externí odkaz:	https://explore.openaire.eu/search/publication?articleId=doi_________::651ed766ea28526718c583500a26dca8 https://doi.org/10.1007/s10772-011-9121-5 Zobrazit plný text záznamu Full text from SpringerLink