- Research Article
268
- 10.1016/s1088-467x(99)00028-1
Mining association rules from quantitative data
- Nov 01, 1999
- Intelligent Data Analysis
- Tzung-Pei Hong
Mining association rules from quantitative data
Machine-learning and data-mining techniques have been developed to turn data into useful task-oriented knowledge. Most algorithms for mining association rules identify relationships among transactions using binary values and find rules at a single-concept level. Transactions with quantitative values and items with hierarchical relationships are, however, commonly seen in real-world applications. This paper proposes a fuzzy multiple-level mining algorithm for extracting knowledge implicit in transactions stored as quantitative values. The proposed algorithm adopts a top-down progressively deepening approach to finding large itemsets. It integrates fuzzy-set concepts, data-mining technologies and multiple-level taxonomy to find fuzzy association rules from transaction data sets. Each item uses only the linguistic term with the maximum cardinality in later mining processes, thus making the number of fuzzy regions to be processed the same as the number of original items. The algorithm therefore focuses on the most important linguistic terms for reduced time complexity.
Mining association rules from quantitative data
Mining association rules from quantitative data
Mining association rules with linguistic terms
Some problems of mining association rules with linguistic terms are discussed. First, an incremental updating algorithm of association rules with linguistic terms is presented. The collection of frequent linguistic attribute sets and its negative border along with their support count are maintained, which makes scan the entire database once at most in the process of updating association rules. The experiment shows that the updating algorithm can not only update association rules effectively but also avoid the repeated cost. Secondly, the parallel algorithm for mining association rules with linguistic terms is presented. The Boolean parallel mining algorithm is improved to discover frequent linguistic attribute sets, and the association rules with at least confidence are generated on all processors. This parallel mining algorithm has fine scale-up, size-up and speed-up.
Read moreAn association rule mining approach for intelligent tutoring system
Intelligent tutoring system (ITS) creates a new teaching mode, but most ITS are merely e-learning platforms that provide course study, without considering learning processes of learners, which can't effectively help learners to consolidate and review the unmastered knowledge points. Data mining techniques can extract the potential, valuable pattern or regulation from a great quantity of data. An intelligent tutoring system has been designed based on data mining technology that could return the learners feedback about knowledge points. In order to quickly find all frequent patterns, i.e., knowledge points, an improved algorithm for mining association rules based on FP-growth is presented. Experimental results show that the improved algorithm can provide effective decision support, and help learners to improve their learning efficiency.
Read morePrivacy Preserving Mining System of Association Rules in OpenStack-Based Cloud
As an efficient data analysis tool, data mining can discover the potential association and regularity of massive data, and it has been widely used and played an important role in business decision, medical research and so on. However, the data mining technology is also a double-edged sword, in bringing convenience at the same time, will also cause the user’s privacy leak problem. In order to solve the problem, the symmetric searchable encryption technology is introduced into the association rule mining system to protect the privacy, and privacy preserving mining system of s (PP-MSAR) in OpenStack-based Cloud environment is designed. In order to solve the problem that the existing data mining algorithms can’t deal with large-scale data, this paper uses the computational power of Hadoop platform and add the global pruning technique to the existing algorithm based on MapReduce association rules, so that the counting of frequent item sets get reduced. At the same time, this paper add frequent matrix storage method into the distributed association rules algorithm and realize he algorithm of mining association rules for frequent matrix storage based on MapReduce. In addition, the introduction of symmetric searchable encryption technology to support the cloud server-side ciphertext retrieval, on the one hand to ensure that users stored in the database information will not be leaked to the outside for others, on the other hand also to ensure that the user data for the system staff confidential. Finally, we test the system, and the results show that the system can carry out association rules mining under the premise of protecting user privacy, and provide the correlation degree between data, which has certain practical significance and application value.
Read moreSufRec, an algorithm for mining association rules: Recursivity and task parallelism
SufRec, an algorithm for mining association rules: Recursivity and task parallelism
A new method to mine valid association rules
To reduce invalid rules in the mining of association rules, we have analyzed the reasons and presented a relative confidence in the judgment criteria. Based on the value of relative confidence, we classify strong association rules into positive, invalid and negative association rules. We offer an algorithm of mining association rules with new judgment criterion and make tests with Visual FoxPro. The tests indicate that the new method stated in this paper can obviously reduce invalid association rules.
Read moreCluster-Based Membership Function Acquisition Approaches for Mining Fuzzy Temporal Association Rules
In real-world applications, transactions are typically represented by quantitative data. Thus, fuzzy association rule mining algorithms have been proposed to handle these quantitative transactions. In addition, items generally have certain lifespans or temporal periods in which they exist in a database. Therefore, fuzzy temporal association rule mining algorithms have also been proposed in the literature. A key factor in the acquisition of fuzzy temporal association rules (FTARs) is the design of appropriate membership functions. Because current approaches have been designed to generate membership functions for mining fuzzy association rules (FARs) in market-basket analysis, in this paper, we propose a membership function tuning mechanism for a fuzzy temporal association rule mining algorithm. The proposed approach modifies an existing cluster-based method to generate unique membership functions that are specifically tailored to each item in a dataset. Two factors are utilized to decide the appropriate membership functions of each item: (1) the density similarity among intervals corresponding to the density similarity within intervals, and (2) the information closeness within an interval corresponding to the similarity in the number of data points between intervals. A parameter θ is used to indicate the relative importance of these two factors. As a result, the membership functions are generated based on the quantitative ranges of individual items, and the generated membership functions of items are different in terms of the values of each interval and the number of intervals. The generated membership functions are subsequently used in a fuzzy temporal association rule mining algorithm. Computational experiments were conducted on both a synthetic dataset and a real-world one to demonstrate the effectiveness of the proposed approach.
Read moreMining Direct and Indirect Weighted Fuzzy Association Rules in Large Transaction Databases
Association rule is an important research topic in data mining and knowledge discovery. Traditional algorithms for mining association rules are built on the binary attributes databases, which has three limitations. Firstly, it can not concern quantitative attributes; secondly, it treats each item with the same significance although different item may have different significance; thirdly, only the direct association rules are discovered. Mining fuzzy association rules has been proposed to address the first limitation. In this paper, we put forward an idea for mining indirect weighted association rules to resolve the other two limitations, and a discovery algorithm for mining both direct and indirect weighted fuzzy association rules by integrating these three extensions.
Read moreMining Positive and Negative Weighted Fuzzy Association Rules in Large Transaction Databases
Association rules mining is an important research topic in data mining and knowledge discovery. Traditional algorithms for mining association rules are built on the binary attributes databases, which has three limitations. Firstly, it cannot concern quantitative attributes; secondly, only the positive association rules are discovered; thirdly, it treat each item with the same significance although different item may have different significance. In this paper, we put forward a discovery algorithm for mining positive and negative fuzzy weighted association rules to resolve these three limitations.
Read moreA parameterised algorithm for mining association rules
A central part of many algorithms for mining association rules in large data sets is a procedure that finds so called frequent itemsets. This paper proposes a new approach to finding frequent itemsets. The approach reduces a number of passes through an input data set and generalises a number of strategies proposed so far. The idea is to analyse a variable number n of itemset lattice levels in p scans through an input data set. It is shown that for certain values of parameters (n,p) this method provides more flexible utilisation of fast access transient memory and faster elimination of itemsets with low support factor. The paper presents the results of experiments conducted to find how the performance of the association rule mining algorithm depends on the values of parameters (n,p).
Read moreComparison and improvement of association rule mining algorithm
In recent years, the data mining technology has been developed rapidly. New efficient algorithms are emerging. Association data mining plays an important role in data mining, and the frequent item sets are the highest and the most costly. This paper is based on the association rules data mining technology. The advantages and disadvantages of Apriori algorithm and FP-growth algorithm are deeply analyzed in the association rules, and a new algorithm is proposed, finally, the performance of the algorithm is compared with the experimental results. It provides a reference for the extension and improvement of the algorithm of association rule mining.
Read moreMining positive and negative fuzzy association rules with multiple minimum supports
Association rules mining is an important research topic in data mining and knowledge discovery. Traditional algorithms for mining association rules are built on the binary attributes databases, which has three limitations. Firstly, it can not concern quantitative attributes; secondly, only the positive association rules are discovered; thirdly, it treat each item with the same frequency although different item may have different frequency. In this paper, we put forward a discovery algorithm for mining positive and negative fuzzy association rules to resolve these three limitations.
Read moreAlgorithm for finding association rules in distributed databases
In the emerging networked environment we are encountering situations in which databases residing at geographically distinct sites must collaborate with each other to analyze their data together. But due to large sizes of the datasets it is neither feasible nor safe to transport large datasets across the network to some common server. We need algorithms that can process the databases at their own locations by exchanging needed information among them and obtain the same results that would have been obtained if the databases were merged. In this paper we present an algorithm for mining association rules from distributed databases by exchanging only the needed summaries among them.
Read moreG3PARM: A Grammar Guided Genetic Programming algorithm for mining association rules
This paper presents the G3PARM algorithm for mining representative association rules. G3PARM is an evolutionary algorithm that uses G3P (Grammar Guided Genetic Programming) and an auxiliary population made up of its best individuals who will then act as parents for the next generation. Due to the nature of G3P, the G3PARM algorithm allows us to obtain valid individuals by defining them through a context-free grammar and, furthermore, this algorithm is generic with respect to data type. We compare our algorithm to two multiobjective algorithms frequently used in literature and known as NSGA2 (Non dominated Sort Genetic Algorithm) and SPEA2 (Strength Pareto Evolutionary Algorithm) and demonstrate the efficiency of our algorithm in terms of running-time, coverage and average support, providing the user with high representative rules.
Read moreImprove efficiency of fuzzy association rule using hedge algebra approach
A major problem when conducting mining fuzzy association rules from the database (DB) is the large computation time and memory needed. In addition, the selection of fuzzy sets for each attribute of the database is very important because it will affect the quality of the mining rule. This paper proposes a method for mining fuzzy association rules using the compressed database. We also use the approach of Hedge Algebra (HA) to build the membership function for attributes instead of using the normal way of fuzzy set theory. This approach allows us to explore fuzzy association rules through a relatively simple algorithm which is faster in terms of time, but it still brings association rules which are as good as the classical algorithms for mining association rules.
Read more