Preventing Early Endpointing for Online Automatic Speech Recognition

Autor:	Shafiq Joty, Yingzhu Zhao, Cheung-Chi Leung, Eng Siong Chng, Bin Ma, Chongjia Ni
Rok vydání:	2021
Předmět:	Ground truth Signal processing Computer science Speech recognition Perspective (graphical) Leverage (statistics) Baseline (configuration management) Security token Training methods Term (time)
Zdroj:	ICASSP
DOI:	10.1109/icassp39728.2021.9413613
Popis:	With the recent development of end-to-end models in speech recognition, there have been more interests in adapting these models for online speech recognition. However, using end-to-end models for online speech recognition is known to suffer from an early endpointing problem, which brings in many deletion errors. In this paper, we propose to address the early endpointing problem from the gradient perspective. Specifically, we leverage on the recently proposed ScaleGrad technique, which was proposed to mitigate the text degeneration issue. Different from ScaleGrad, we adapt it to discourage the early generation of the end-of-sentence ( ) token. A scaling term is added to directly maneuver the gradient of the training loss to encourage the model to learn to keep generating non- tokens. Compared with previous approaches such as voice-activity-detection and end-of-query detection, the proposed method does not rely on various types of silence, and it also saves the trouble from obtaining the ground truth endpoint with forced alignment. Nevertheless, it can be jointly applied with other techniques. Experiments on AISHELL-1 dataset show that our model brings relative 5.4%-10.1% CER reductions over the baseline, and surpasses the unlikelihood training method which directly reduces the generation probability of token.
Databáze:	OpenAIRE
Externí odkaz:	https://explore.openaire.eu/search/publication?articleId=doi_________::0157bb4a16a85dc76fd042f39082b947 https://doi.org/10.1109/icassp39728.2021.9413613 Zobrazit plný text záznamu