Enhanced Deep Speech Separation in Clustered Ad Hoc Distributed Microphone Environments

Autor:	Kim, Jihyun, Kindt, Stijn, Madhu, Nilesh, Kang, Hong-Goo
Rok vydání:	2024
Předmět:	Electrical Engineering and Systems Science - Audio and Speech Processing
Druh dokumentu:	Working Paper
Popis:	Ad-hoc distributed microphone environments, where microphone locations and numbers are unpredictable, present a challenge to traditional deep learning models, which typically require fixed architectures. To tailor deep learning models to accommodate arbitrary array configurations, the Transform-Average-Concatenate (TAC) layer was previously introduced. In this work, we integrate TAC layers with dual-path transformers for speech separation from two simultaneous talkers in realistic settings. However, the distributed nature makes it hard to fuse information across microphones efficiently. Therefore, we explore the efficacy of blindly clustering microphones around sources of interest prior to enhancement. Experimental results show that this deep cluster-informed approach significantly improves the system's capacity to cope with the inherent variability observed in ad-hoc distributed microphone environments. Comment: Accepted to Interspeech 2024
Databáze:	arXiv
Externí odkaz:	http://arxiv.org/abs/2406.09819 Zobrazit plný text záznamu View this record from Arxiv