Toward human-assisted lexical unit discovery without text resources
Autor: | Chris Bartels, Colleen Richey, Chiachi Hung, Andreas Kathol, Vikramjit Mitra, Dimitra Vergyri, Harry Bratt, Wen Wang |
---|---|
Rok vydání: | 2016 |
Předmět: |
Matching (statistics)
Computer science business.industry String (computer science) 02 engineering and technology Pragmatics USable computer.software_genre Fuzzy logic Lexical item Data modeling 03 medical and health sciences 0302 clinical medicine Phone 030221 ophthalmology & optometry 0202 electrical engineering electronic engineering information engineering 020201 artificial intelligence & image processing Artificial intelligence business computer Natural language processing |
Zdroj: | SLT |
DOI: | 10.1109/slt.2016.7846246 |
Popis: | This work addresses lexical unit discovery for languages without (usable) written resources. Previous work has addressed this problem using entirely unsupervised methodologies. Our approach in contrast investigates the use of linguistic and speaker knowledge which are often available even if text resources are not. We create a framework that benefits from such resources, not assuming orthographic representations and avoiding generation of word-level transcriptions. We adapt a universal phone recognizer to the target language and use it to convert audio into a searchable phone string for lexical unit discovery via fuzzy sub-string matching. Linguistic knowledge is used to constrain phone recognition output and to constrain lexical unit discovery on the phone recognizer output. |
Databáze: | OpenAIRE |
Externí odkaz: |