Database Similarity Join for Metric Spaces

Autor: Jason A. Cheney, Spencer S. Pearson, Yasin N. Silva
Rok vydání: 2013
Předmět:
Zdroj: Similarity Search and Applications ISBN: 9783642410611
SISAP
DOI: 10.1007/978-3-642-41062-8_27
Popis: Similarity Joins are recognized among the most useful data processing and analysis operations. They retrieve all data pairs whose distances are smaller than a predefined threshold e. While several standalone implementations have been proposed, very little work has addressed the implementation of Similarity Join as a physical database operator. In this paper, we focus on the study, design and implementation of a Similarity Join database operator for any dataset that lies in a metric space DBSimJoin. We describe the changes in each query engine module to implement DBSimJoin and provide details of our implementation in PostgreSQL. The extensive performance evaluation shows that DBSimJoin significantly outperforms alternative approaches.
Databáze: OpenAIRE