Cost-Based and Effective Human-Machine Based Data Deduplication Model in Entity Reconciliation
Autor: | Lawrence Tandoh, MengShu Hou, Barbie Eghan-Yartel, Maame G. Asante-Mensah, Moses J. Eghan, Michael Y. Kpiebaareh, Charles Roland Haruna |
---|---|
Rok vydání: | 2018 |
Předmět: |
Computer science
business.industry media_common.quotation_subject 02 engineering and technology computer.software_genre Crowdsourcing Data set 020204 information systems 0202 electrical engineering electronic engineering information engineering Benchmark (computing) Data deduplication 020201 artificial intelligence & image processing Human–machine system Quality (business) Data mining Cluster analysis business computer media_common |
Zdroj: | ICSAI |
DOI: | 10.1109/icsai.2018.8599375 |
Popis: | In real world, databases often have several records representing the same entity and these duplicates have no common key, thus making deduplication difficult. Machine-based and crowdsourcing techniques were disjointly used in improving quality in data deduplication. Crowdsourcing were used for solving tasks that the machine-based algorithms were not good at. Though, the crowds, compared with machines, provided relatively more accurate results, both platforms were slow in execution and hence expensive to implement. In this paper, a hybrid human-machine system was proposed where machines were firstly used on the data set before the humans were further used to identify potential duplicates. We performed experiments using three benchmark datasets; paper, restaurant and product datasets. Our algorithm was compared with some existing techniques and our approach outperformed some methods by achieving a high accuracy of deduplication and good deduplication efficiency while incurring low crowdsourcing costs. |
Databáze: | OpenAIRE |
Externí odkaz: |