Modeling prosodic dynamics for speaker recognition

Autor:	R. Mihaescu, John J. Godfrey, André Gustavo Adami, Douglas A. Reynolds
Rok vydání:	2004
Předmět:	Dynamic time warping Computer science business.industry Speech recognition Bigram Feature extraction Word error rate Pattern recognition Fundamental frequency Speech processing Speaker recognition Speaker diarisation NIST Artificial intelligence business
Zdroj:	ICASSP (4)
DOI:	10.1109/icassp.2003.1202761
Popis:	Most current state-of-the-art automatic speaker recognition systems extract speaker-dependent features by looking at short-term spectral information. This approach ignores long-term information that can convey supra-segmental information, such as prosodics and speaking style. We propose two approaches that use the fundamental frequency and energy trajectories to capture long-term information. The first approach uses bigram models to model the dynamics of the fundamental frequency and energy trajectories for each speaker. The second approach uses the fundamental frequency trajectories of a predefined set of words as the speaker templates and then, using dynamic time warping, computes the distance between the templates and the words from the test message. The results presented in this work are on Switchboard I using the NIST Extended Data evaluation design. We show that these approaches can achieve an equal error rate of 3.7%, which is a 77% relative improvement over a system based on short-term pitch and energy features alone.
Databáze:	OpenAIRE
Externí odkaz:	https://explore.openaire.eu/search/publication?articleId=doi_________::1db926d22dad1bb0eb1f85713cc0b833 https://doi.org/10.1109/icassp.2003.1202761 Zobrazit plný text záznamu