Reliable representations for association rules
Autor: | Yue Xu, Yuefeng Li, Gavin Shaw |
---|---|
Rok vydání: | 2011 |
Předmět: |
Association rule mining
Information Systems and Management Association rule learning Basis (linear algebra) Computer science business.industry media_common.quotation_subject Redundant Association rules Inference computer.software_genre Machine learning Closed itemsets Set (abstract data type) Knowledge discovery Knowledge extraction Redundancy (engineering) Quality (business) Artificial intelligence Data mining business Representation (mathematics) computer 080600 INFORMATION SYSTEMS media_common |
Zdroj: | Data & Knowledge Engineering |
ISSN: | 2381-3652 |
Popis: | Association rule mining has contributed to many advances in the area of knowledge discovery. However, the quality of the discovered association rules is a big concern and has drawn more and more attention recently. One problem with the quality of the discovered association rules is the huge size of the extracted rule set. Often for a dataset, a huge number of rules can be extracted, but many of them can be redundant to other rules and thus useless in practice. Mining non-redundant rules is a promising approach to solve this problem. In this paper, we first propose a definition for redundancy, then propose a concise representation, called a Reliable basis, for representing non-redundant association rules. The Reliable basis contains a set of non-redundant rules which are derived using frequent closed itemsets and their generators instead of using frequent itemsets that are usually used by traditional association rule mining approaches. An important contribution of this paper is that we propose to use the certainty factor as the criterion to measure the strength of the discovered association rules. Using this criterion, we can ensure the elimination of as many redundant rules as possible without reducing the inference capacity of the remaining extracted non-redundant rules. We prove that the redundancy elimination, based on the proposed Reliable basis, does not reduce the strength of belief in the extracted rules. We also prove that all association rules, their supports and confidences, can be retrieved from the Reliable basis without accessing the dataset. Therefore the Reliable basis is a lossless representation of association rules. Experimental results show that the proposed Reliable basis can significantly reduce the number of extracted rules. We also conduct experiments on the application of association rules to the area of product recommendation. The experimental results show that the non-redundant association rules extracted using the proposed method retain the same inference capacity as the entire rule set. This result indicates that using non-redundant rules only is sufficient to solve real problems needless using the entire rule set. |
Databáze: | OpenAIRE |
Externí odkaz: |