- Research Article
59
- 10.1016/j.patcog.2016.12.003
Ground truth bias in external cluster validity indices
- Dec 08, 2016
- Pattern Recognition
- Yang Lei + 5 more +5
Ground truth bias in external cluster validity indices
The existence of large volumes of time series data in many applications has motivated data miners to investigate specialized methods for mining time series data. Clustering is a popular data mining method due to its powerful exploratory nature and its usefulness as a preprocessing step for other data mining techniques. This article develops two novel clustering algorithms for time series data that are extensions of a crisp c-shapes algorithm. The two new algorithms are heuristic derivatives of fuzzy c-means (FCM). Fuzzy c-Shapes plus (FCS+) replaces the inner product norm in the FCM model with a shape-based distance function. Fuzzy c-Shapes double plus (FCS++) uses the shape-based distance, and also replaces the FCM cluster centers with shape-extracted prototypes. Numerical experiments on 48 real time series data sets show that the two new algorithms outperform state-of-the-art shape-based clustering algorithms in terms of accuracy and efficiency. Four external cluster validity indices (the Rand index, Adjusted Rand Index, Variation of Information, and Normalized Mutual Information) are used to match candidate partitions generated by each of the studied algorithms. All four indices agree that for these finite waveform data sets, FCS++ gives a small improvement over FCS+, and in turn, FCS+ is better than the original crisp c-shapes method. Finally, we apply two tests of statistical significance to the three algorithms. The Wilcoxon and Friedman statistics both rank the three algorithms in exactly the same way as the four cluster validity indices.
Ground truth bias in external cluster validity indices
Ground truth bias in external cluster validity indices
A Comparison Between External and Internal Cluster Validity Indices
This paper presents a comparison between external and internal cluster validity indices with a similar bounded index range. F-measure (FM) and Fowlkes-Mallows (FMI) of external validity indices, as well as Silhouette (SIL) of internal validity index, were chosen for this comparative analysis. Ten numerical data sets, namely Haberman, BUPA (liver disorder), Wisconsin Diagnostic Breast Cancer (WDBC), Iris, Seeds, Wine, User Knowledge, Cleveland, Segmentation, and Glass, were deployed to benchmark the clustering outcomes based on Fuzzy C-Mean (FCM) algorithm. Mean, minimum, and maximum scores were calculated to determine the similarities and differences among the indices. Pearson correlation was reported as well. As a result, the index scores displayed a slight difference between the external and internal validity indices. A moderate and positive correlation was noted between the external and internal validity index (r=. 66, r=.65 p<0.01) scores. This correlation signifies a similar graph pattern between the cluster validity indices. This comparative analysis revealed that the external and internal cluster validity indices with similar bounded index ranges and slightly different index scores generate a moderate and positive correlation with a similar graph pattern.
Read moreSimrec: a similarity measure recommendation system for mixed data clustering algorithms
Clustering algorithms play a pivotal role in data mining, offering powerful tools for uncovering hidden patterns and structures within datasets. These algorithms aim to divide data points into coherent groups based on similarities or dissimilarities, making it easier to explore and understand complex data. Clustering algorithms typically rely on similarity measures to assess the likeness between data points. Consequently, selecting a suitable similarity measure is crucial for achieving satisfactory clustering outcomes. However, this decision can pose significant challenges, especially for non-experts, given the plethora of similarity measures available in the literature and their performance which is closely linked to the specific dataset, clustering algorithm, and cluster validity index employed. This difficulty is even more important when considering mixed data clustering. Mixed data refers to heterogeneous data characterized by both numerical and categorical attributes. In such a context, the same similarity measure cannot be used for both types of attributes due to their different nature. Commonly, two similarity measures are combined, one for numerical attributes and one for categorical attributes. This adds a layer of complexity to the problem since it requires the selection of two similarity measures instead of just one. This paper introduces SIMREC, a similarity measure recommendation system for mixed data clustering. The system uses meta-learning to mine the relationship between dataset characteristics and similarity measures performances for different mixed data clustering algorithms and cluster validity indices. Therefore, given a mixed dataset, a mixed data clustering algorithm, and a cluster validity index, the system can recommend suitable pairs of numerical and categorical similarity measures based on the characteristics of the dataset. We implemented the proposed system using 130 pairs of similarity measures (10 numerical and 13 categorical), 4 commonly used mixed data clustering algorithms (K-Prototypes, LSH-K-Prototypes, K-Medoids, and Hierarchical Clustering), and three cluster validity indices (Silhouette, Clustering Accuracy, and Adjusted Rand Index). Our experiments on 185 publicly available mixed datasets show that the pairs of similarity measures recommended by SIMREC outperform the baseline pairs, including classically used pairs of similarity measures in the literature.
Read moreTS-Benchmark: A Benchmark for Time Series Databases
Time series data is widely used in scenarios such as supply chain, stock data analysis, and smart manufacturing. A number of time series database systems have been invented to manage and query large volumes of time series data. We observe that the existing benchmarks of time series databases are focused on workloads of complex analysis such as pattern matching and trend prediction whose performance may be highly affected by the data analysis algorithms, instead of the back-end databases. However, in many real applications of time series databases, people are more interested in the performance metrics such as data injection throughput and query processing time. A benchmark is still required to extensively compare the performance of time series databases in such metrics. We introduce such a benchmark called TS-Benchmark which majorly applies a scenario of device monitoring for wind turbines. A DCGAN-based data generation model is proposed to generate large volumes of time series data from some real time series data. The workloads are categorized into three folds: data loading (in batch), streaming data injection, and historical data access (for typical queries). We implement the benchmark and compare four representative time series databases: InfluxDB, TimescaleDB, Druid and OpenTSDB. The results are reported and analyzed.
Read moreUse of a fuzzy granulation–degranulation criterion for assessing cluster validity
Use of a fuzzy granulation–degranulation criterion for assessing cluster validity
IMI2: A fuzzy clustering validity index for multiple imbalanced clusters
IMI2: A fuzzy clustering validity index for multiple imbalanced clusters
Fuzzy c-means clustering algorithm with unknown number of clusters for symbolic interval data
In this study, the concepts of competitive agglomeration clustering algorithm is incorporated into fuzzy c-means (FCM) clustering algorithm for symbolic interval-values data. In the proposed approach, called as IFCMwUNC clustering algorithm, the problems of the unknown clusters number and the initialization of prototypes in the FCM clustering algorithm for symbolic interval-values data are overcome and discussed. Due to the competitive agglomeration clustering algorithm possess the advantages of the hierarchical clustering algorithm and the partitional clustering algorithm, IFCMwUNC clustering algorithm can be fast converges in a few iterations regardless of the initial number of clusters. Moreover, it is also converges to the same optimal partition regardless of its initialization. Experiments results show the merits and usefulness of IFCMwUNC clustering algorithm for the symbolic interval-values data.
Read moreUsing HBase to Implement Speed Layer in Time Series Data Storage Systems
In recent years, modern systems have become increasingly integrated, and the challenges are focused on delivering real-time analytics based on big data. Thus, using standard software tools to extract information from such datasets is not always possible. The Lambda Architecture proposed by Marz is an architectural solution that can manage the processing of large data volumes by combining real-time and data batch processing techniques. Choosing a suitable database management system for storing large volumes of time series data is not a trivial issue as various aspects such as low latency, high performance and the possibility of horizontal scalability must be taken into account. The new NoSQL approaches use for this purpose non-relational databases with significant advantages in terms of flexibility and performance in comparison with the traditional relational databases. With reference to this, the purpose of this paper is to analyse the general characteristics of time series data and the main activities performed by the Speed layer in a system based on the Lambda Architecture. Based on this, the use of a column-oriented NoSQL DBMS as a system for storing time series data is justified. The paper also addresses the challenges of using HBase as a system for storing and analysing time series data. These questions are related to the design of an appropriate database schema, the need to achieve balance between ease of access to the data and performance as well as considering the factors that affect the overload of individual nodes in the system.
Read moreReclust: an efficient clustering algorithm for mixed data based on reclustering and cluster validation
<span>Clustering is a significant approach in data mining, which seeks to find groups or clusters of data. Both numeric and categorical features are frequently used to define the data in real-world applications. Several different clustering algorithms are proposed for the numerical and categorical datasets. In clustering algorithms, the quality of clustering results is evaluated using cluster validation. This paper proposes an efficient clustering algorithm for mixed numerical and categorical data using re-clustering and cluster validation. Initially, the mixed dataset is clustered with four traditional clustering algorithms like expectation-maximization (EM), hierarchical cluster (HC), k-means (KM), and self-organizing map (SOM). These four algorithms are validated, and the best algorithm is selected for re-clustering. It is an iterative process for improving the quality of cluster results. The incorrectly clustered data is iteratively re-clustered and evaluated based on the cluster validation. The performance of the proposed clustering method is evaluated with a real-time dataset in terms of purity, normalized mutual information, rand index, precision, and recall. The experimental results have shown that the proposed reclust algorithm achieves better performance compared to other clustering algorithms.</span>
Read moreObjective evaluation of knitted yarn quality based on fuzzy kernel C-means
In this paper, the method that measuring dataset of knitted yarns is clustered using improving fuzzy kernel c-Means (FKCM) clustering algorithm is proposed. In FKCM clustering algorithm, the data of low dimension input space is mapped to high dimension feature space, FCM clustering algorithm is performed in feature space, then the constraint optimization distance matrix and membership matrix of testing samples are computed by utilizing iterative algorithm, and the clustering result can be acquired according maximum membership principle. Subsequently, the Kernel F cluster validity index is designed for seeking the fitness cluster number and the corresponding relationship model of clusters sequence number and quality grades is constructed. Improving FKCM clustering algorithm is more efficient than other clustering algorithms, and clustering result can provide the training samples for constructing quality grades and clusters recognition function of new samples. The combination of improving FKCM and KF index provides an efficient data analysis method for multi-index dataset.
Read moreMulti-objective clustering of tissue samples for cancer diagnosis
In the field of pattern recognition, the study of the gene expression profiles for different tissue samples over different experimental conditions has became feasible with the arrival of micro-array based technology. In cancer research, classification of tissue samples is necessary for cancer diagnosis, which can be done with the help of micro-array technology. In this article we have presented a multi-objective optimization ( MOO ) based clustering technique utilizing AMOSA ( Archived Multi-Objective Simulated Annealing ) as the underlying optimization strategy for classification of tissue samples from cancer data sets. As objective functions three cluster validity indices namely, XB, PBM, and FCM indices are optimized simultaneously to form more accurate clusters of tissue samples. The presented clustering technique is evaluated for two open source benchmark cancer data sets, which are Brain tumor data set and Adult Malignancy data set. In order to evaluate the quality or goodness of produced clusters two cluster quality measures viz, Adjusted Rand Index ( ARI ) and Classification Accuracy ( %CoA ) are calculated for each data set. Comparative results of the presented clustering algorithm with 10 state-of-the-art existing single-objective, multi-objective based clustering algorithms are shown for two benchmark data sets.
Read moreAutomatic machine embroidery image color analysis system. Part I: Using Gustafson-Kessel clustering algorithm in embroidery fabric color separation
This series of studies aims to propose the automatic machine embroidery image color analysis system to solve the problem of lack of manpower for machine embroidery fabric drafting and to shorten drafting time. The studies included three parts: (1) machine embroidery image color separation, (2) search of repeated pattern images, (3) machine embroidery color analysis system integration. This study aimed to find the optimal clustering algorithm and cluster validity indices for the automatic color separation process of machine embroidery fabric drafting in order to shorten drafting time. To improve image quality for computer analysis, this study used the color hybrid median filter to filter noise and the color bilateral filter to smoothen fabric and embroidery texture for subsequent color separation. By extracting the color a* component and b* component of the machine embroidery image in CIE L*a*b* color system, this study used the Gustafson-Kessel clustering algorithm for color separation. The Gustafson-Kessel clustering algorithm in machine embroidery image color separation can improve color separation accuracy, and its result is compared with that of the clustering algorithms commonly used in the color separation of color images. This study implemented the chromatography of the color separation results, and used the cluster validity indices to prove that the application of Gustafson-Kessel clustering algorithm in the machine embroidery image color separation system has better results than K-means, K-medoid, fuzzy C-means (FCM), and self-organizing map (SOM) clustering algorithm. The results meet the classifications as expected by human eyes.
Read moreOptimized Fuzzy Clustering Algorithms for Brain MRI Image Segmentation Based on Local Gaussian Probability and Anisotropic Weight Models
Brain Magnetic Resonance Imaging (MRI) image segmentation is one of the critical technologies of clinical medicine, and is the basis of three-dimensional reconstruction and downstream analysis between normal tissues and diseased tissues. However, there are various limitations in brain MRI images, such as gray irregularities, noise, and low contrast, reducing the accuracy of the brain MRI images segmentation. In this paper, we propose two optimization solutions for the fuzzy clustering algorithm based on local Gaussian probability fuzzy C-means (LGP-FCM) model and anisotropic weight fuzzy C-means (AW-FCM) model and apply it in brain MRI image segmentation. An FCM clustering algorithm is proposed based on AW-FCM. By introducing the new neighborhood weight calculation method, each point has the weight of anisotropy, effectively overcomes the influence of noise on the image segmentation. In addition, the LGP model is introduced in the objective function of fuzzy clustering, and a fuzzy clustering segmentation algorithm based on LGP-FCM is proposed. A clustering segmentation algorithm of adaptive scale fuzzy LGP model is proposed. The neighborhood scale corresponding to each pixel in the image is automatically estimated, which improves the robustness of the model and achieves the purpose of precise segmentation. Extensive experimental results demonstrate that the proposed LGP-FCM algorithm outperforms comparison algorithms in terms of sensitivity, specificity and accuracy. LGP-FCM can effectively segment the target regions from brain MRI images.
Read moreBig Data Technology for Comparative Study of K-Means and Fuzzy C-Means Algorithms Performance
Big data is technology that has the ability to manage very large amounts of data, in very fast time to allow real-time analysis and reactions. Several clustering methods which are used to group data are Fuzzy C-Means (FCM) and K-Means Clustering. K-Means Clustering algorithm is a method of partitioning existing data into two or more group. This research goal was to compare the performance of K-Means and Fuzzy C-Means algorithms in clustering data using big data technology. In this research, Hadoop and Hive were chosen the big data technology. The knowledge of Shia history on student and lecturer of Syarif Hidayatullah State Islamic University Jakarta were the data which used in this research, The testing was done by constructing application K-Means Fuzzy C-Means using Java language, Hadoop and Hive and then test the performance of K-Means and Fuzzy C-Means algorithms in data clustering. It compares both algorithms in terms of accuracy, execution time, and time complexity of the algorithms. In the application K-Means Fuzzy C-Means, evaluation were performed with data filter and the average accuracy difference result of K-Means and Fuzzy C-Means is 8.03% with the better accuracy owned by K-Means. The average execution time difference is 718.58 ms, which K-Means was faster than is Fuzzy C-Means. The time complexities of both algorithms have the same value O(n2) and the Big O equation resulted in an average difference of 93,568 with the smallest value on K-Means. Thus, K-Means algorithm is better than the Fuzzy C-Means in terms of accuracy, execution time, and the time complexity.
Read moreCorrelation Based Cluster Validity Index for Recognition of Leukemia Mediating Biomarkers
This article proposes an index, known as Correlation Index (CRI) with respect to correlation coefficient, to validate the clusters from a clustering algorithm. The index is utilized to recognize some genomes that have been changed quite subsequently from normal state to disease state according to their expression impressions. Firstly, a clustering technique has been adapted on microarray dataset to generate the number of clusters. Thereafter, proposed correlation index is applied on distributed clusters to calculate correlation between two genes and select altered gene sets. The CRI based cluster validity index has been exercised on Leukemia microarray dataset for the recognition of biomarkers which are significantly depicted from normal phase to cancerous phase. In this connection, we have validated the altered gene set using the gene ontology (GO) attribute based on p-value statistics. Therefore, some disease mediating biomarkers have been recognized from gene expression dataset. Here, three clustering algorithms, viz., k-means, PAM, and Fuzzy c-means have been applied on the datasets. The proposed CRI method achieves good efficiency in comparison with state-of-the-art cluster validity indices, and the results are appropriately validated using F1-score and top-K gene set of NCBI.
Read more