Why are these similar? Investigating item similarity types in a large digital library
Autor: | Nikolaos Aletras, German Rigau, Mark Stevenson, Eneko Agirre, Aitor Gonzalez-Agirre |
---|---|
Rok vydání: | 2015 |
Předmět: |
Information Systems and Management
Information retrieval Computer Networks and Communications business.industry Computer science Context (language use) 02 engineering and technology Library and Information Sciences Crowdsourcing Digital library Pearson product-moment correlation coefficient Personalization Task (project management) Set (abstract data type) symbols.namesake Similarity (network science) 020204 information systems 0202 electrical engineering electronic engineering information engineering symbols 020201 artificial intelligence & image processing business Information Systems |
Zdroj: | Journal of the Association for Information Science and Technology. 67:1624-1638 |
ISSN: | 2330-1635 |
DOI: | 10.1002/asi.23482 |
Popis: | We introduce a new problem, identifying the type of relation that holds between a pair of similar items in a digital library. Being able to provide a reason why items are similar has applications in recommendation, personalization, and search. We investigate the problem within the context of Europeana, a large digital library containing items related to cultural heritage. A range of types of similarity in this collection were identified. A set of 1,500 pairs of items from the collection were annotated using crowdsourcing. A high intertagger agreement average 71.5 Pearson correlation was obtained and demonstrates that the task is well defined. We also present several approaches to automatically identifying the type of similarity. The best system applies linear regression and achieves a mean Pearson correlation of 71.3, close to human performance. The problem formulation and data set described here were used in a public evaluation exercise, the *SEM shared task on Semantic Textual Similarity. The task attracted the participation of 6 teams, who submitted 14 system runs. All annotations, evaluation scripts, and system runs are freely available. |
Databáze: | OpenAIRE |
Externí odkaz: |