Multichannel Overlapping Speaker Segmentation Using Multiple Hypothesis Tracking Of Acoustic And Spatial Features
Autor: | Patrick A. Naylor, Aidan O. T. Hogg, Christine Evers |
---|---|
Přispěvatelé: | Engineering and Physical Sciences Research Council |
Rok vydání: | 2021 |
Předmět: |
Technology
speaker segmentation multiple hypothesis tracking Computer science Speech recognition direction of arrival Computer Science Artificial Intelligence Engineering Segmentation fundamental frequency Imaging Science & Photographic Technology Signal processing Science & Technology business.industry Deep learning Search engine indexing Direction of arrival Engineering Electrical & Electronic Acoustics Computer Science Software Engineering Speaker diarisation Computer Science Hit rate DIARIZATION Kalman filter Artificial intelligence business Focus (optics) |
Zdroj: | ICASSP IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) |
Popis: | An essential part of any diarization system is the task of speaker segmentation which is important for many applications including speaker indexing and automatic speech recognition (ASR) in multi-speaker environments. Segmentation of overlapping speech has recently been a key focus of this work. In this paper we explore the use of a new multimodal approach for overlapping speaker segmentation that tracks both the fundamental frequency (F 0 ) of the speaker and the speaker’s direction of arrival (DOA) simultaneously. Our proposed multiple hypothesis tracking system, which simultaneously tracks both features, shows an improvement in segmentation performance when compared to tracking these features separately. An illustrative example of overlapping speech demonstrates the effectiveness of our proposed system. We also undertake a statistical analysis on 12 meetings from the AMI corpus and show an improvement in the HIT rate of 14.1% on average against a commonly used deep learning bidirectional long short term memory network (BLSTM) approach. |
Databáze: | OpenAIRE |
Externí odkaz: |