SoundCollage: Automated Discovery of New Classes in Audio Datasets

Autor:	Choi, Ryuhaerang, Chatterjee, Soumyajit, Spathis, Dimitris, Lee, Sung-Ju, Kawsar, Fahim, Malekzadeh, Mohammad
Rok vydání:	2024
Předmět:	Computer Science - Sound Electrical Engineering and Systems Science - Audio and Speech Processing
Druh dokumentu:	Working Paper
Popis:	Developing new machine learning applications often requires the collection of new datasets. However, existing datasets may already contain relevant information to train models for new purposes. We propose SoundCollage: a framework to discover new classes within audio datasets by incorporating (1) an audio pre-processing pipeline to decompose different sounds in audio samples and (2) an automated model-based annotation mechanism to identify the discovered classes. Furthermore, we introduce clarity measure to assess the coherence of the discovered classes for better training new downstream applications. Our evaluations show that the accuracy of downstream audio classifiers within discovered class samples and held-out datasets improves over the baseline by up to 34.7% and 4.5%, respectively, highlighting the potential of SoundCollage in making datasets reusable by labeling with newly discovered classes. To encourage further research in this area, we open-source our code at https://github.com/nokia-bell-labs/audio-class-discovery. Comment: 5 pages, 2 figures
Databáze:	arXiv
Externí odkaz:	http://arxiv.org/abs/2410.23008 Zobrazit plný text záznamu View this record from Arxiv