Improved Modeling of Cross-Decoder Phone Co-Occurrences in SVM-Based Phonotactic Language Recognition

Autor:	Mikel Peñagarikano, Germán Bordel, Luis Javier Rodríguez-Fuentes, Amparo Varona
Rok vydání:	2011
Předmět:	Phonotactics Acoustics and Ultrasonics Computer science business.industry Speech recognition Feature extraction Speaker recognition computer.software_genre Support vector machine Phone NIST Artificial intelligence Electrical and Electronic Engineering business Baseline (configuration management) computer Decoding methods Natural language processing
Zdroj:	IEEE Transactions on Audio, Speech, and Language Processing. 19:2348-2363
ISSN:	1558-7924 1558-7916
Popis:	Most common approaches to phonotactic language recognition deal with several independent phone decodings. These decodings are processed and scored in a fully uncoupled way, their time alignment (and the information that may be extracted from it) being completely lost. Recently, we have presented two new approaches to phonotactic language recognition which take into account time alignment information, by considering time-synchronous cross-decoder phone co-occurrences. Experiments on the 2007 NIST LRE database demonstrated that using phone co-occurrence statistics could improve the performance of baseline phonotactic recognizers. In this paper, approaches based on time-synchronous cross-decoder phone co-occurrences are further developed and evaluated with regard to a baseline SVM-based phonotactic system, by using: 1) counts of n-grams (up to 4-grams) of phone co-occurrences; and 2) the degree of co-occurrence of phone n-grams (up to 4-grams). To evaluate these approaches, a choice of open software (Brno University of Technology phone decoders, LIBLINEAR and FoCal) was used, and experiments were carried out on the 2007 NIST LRE database. The two approaches presented in this paper outperformed the baseline phonotactic system, yielding around 7% relative improvement in terms of CLLR. The fusion of the baseline system with the two proposed approaches yielded 1.83% EER and CLLR=0.270 (meaning 18% relative improvement), the same performance (on the same task) than state-of-the-art phonotactic systems which apply more complex models and techniques, thus supporting the use of cross-decoder dependencies for language recognition.
Databáze:	OpenAIRE
Externí odkaz:	https://explore.openaire.eu/search/publication?articleId=doi_________::10b2b4d56ae20165e06e0722167d8715 https://doi.org/10.1109/tasl.2011.2134088 Zobrazit plný text záznamu