The Thinkit System for Icassp2021 M2voc Challenge

Autor:	Zengqiang Shang, Pengyuan Zhang, Haozhe Zhang, Bolin Zhou, Ziyi Chen
Rok vydání:	2021
Předmět:	Set (abstract data type) Signal processing Cloning (programming) Computer science Model selection Speech recognition Boundary (topology) Waveform Acoustic model Prosody
Zdroj:	ICASSP
DOI:	10.1109/icassp39728.2021.9413669
Popis:	In this paper, we introduce the low resource text-to-speech system from the ThinkIT team submitted to Multi-Speaker Multi-Style Voice Cloning Challenge (M2VoC). The challenge has two tasks: few-shot track1 provides 100 samples for each person and one-shot track2 offers 5 samples only. Each track contains two sub-tracks A and B. Instead of sub-track A, sub-track B can use extra public data besides the released data. But we participate in the sub-track A only. We choose the finetune as our backbone strategy. Our submitted systems include BERT based prosody boundary prediction module, FastSpeech based acoustic model to generate acoustic features from text input, and HIFIGAN based vocoder to generate waveform from acoustic features. Among them, acoustic models are susceptible to low resource speakers. To prevent over-fitting, we modified the acoustic model and split out validation set to assist the manual model selection. Evaluation results provided by the challenges organizers demonstrate the effectiveness of our system.
Databáze:	OpenAIRE
Externí odkaz:	https://explore.openaire.eu/search/publication?articleId=doi_________::5b62648161a6051a11c1da0ad69b97d2 https://doi.org/10.1109/icassp39728.2021.9413669 Zobrazit plný text záznamu