Audio-Visual Automatic Speech Recognition Using PZM, MFCC and Statistical Analysis

Autor:	Saswati Debnath, Pinki Roy
Jazyk:	angličtina
Rok vydání:	2021
Předmět:	audio-visual speech recognition lip tracking pseudo zernike moment mel frequency cepstral coefficients (mfcc) incremental feature selection (ifs) statistical analysis Technology
Zdroj:	International Journal of Interactive Multimedia and Artificial Intelligence, Vol 7, Iss 2, Pp 121-133 (2021)
Druh dokumentu:	article
ISSN:	1989-1660 48974714
DOI:	10.9781/ijimai.2021.09.001
Popis:	Audio-Visual Automatic Speech Recognition (AV-ASR) has become the most promising research area when the audio signal gets corrupted by noise. The main objective of this paper is to select the important and discriminative audio and visual speech features to recognize audio-visual speech. This paper proposes Pseudo Zernike Moment (PZM) and feature selection method for audio-visual speech recognition. Visual information is captured from the lip contour and computes the moments for lip reading. We have extracted 19th order of Mel Frequency Cepstral Coefficients (MFCC) as speech features from audio. Since all the 19 speech features are not equally important, therefore, feature selection algorithms are used to select the most efficient features. The various statistical algorithm such as Analysis of Variance (ANOVA), Kruskal-wallis, and Friedman test are employed to analyze the significance of features along with Incremental Feature Selection (IFS) technique. Statistical analysis is used to analyze the statistical significance of the speech features and after that IFS is used to select the speech feature subset. Furthermore, multiclass Support Vector Machine (SVM), Artificial Neural Network (ANN) and Naive Bayes (NB) machine learning techniques are used to recognize the speech for both the audio and visual modalities. Based on the recognition rate combined decision is taken from the two individual recognition systems. This paper compares the result achieved by the proposed model and the existing model for both audio and visual speech recognition. Zernike Moment (ZM) is compared with PZM and shows that our proposed model using PZM extracts better discriminative features for visual speech recognition. This study also proves that audio feature selection using statistical analysis outperforms methods without any feature selection technique.
Databáze:	Directory of Open Access Journals
Externí odkaz:	https://doaj.org/article/57ccb70bfa484c8da2a76e0e48974714 Zobrazit plný text záznamu View record in DOAJ