- Research Article
17
- 10.1016/j.jocs.2022.101925
A varied density-based clustering algorithm
- Dec 12, 2022
- Journal of Computational Science
- Ahmed Fahim
A varied density-based clustering algorithm
Stream data applications have become more and more prominent recently and the requirements for stream clustering algorithms have increased drastically. Due to continuously evolving nature of the stream, it is crucial that the algorithm autonomously detects clusters of arbitrary shape, with different densities, and varying number of clusters. Although available density-based stream clustering are able to detect clusters with arbitrary shapes and varying numbers, they fail to adapt their thresholds to detect clusters with different densities. In this paper we propose a stream clustering algorithm called HASTREAM, which is based on a hierarchical density-based clustering model that automatically detects clusters of different densities. The density thresholds are independently adapted to the existing data without the need of any user intervention. To reduce the high computational cost of the presented approach, techniques from the graph theory domain are utilized to devise an incremental update of the underlying model. To show the effectiveness of HASTREAM and hierarchical density-based approaches in general, several synthetic and real world data sets are evaluated using various quality measures. The results showed that the hierarchical property of the model was able to improve the quality of density-based stream clusterings and enabled HASTREAM to detect streaming clusters of different densities.
A varied density-based clustering algorithm
A varied density-based clustering algorithm
An Extended DBSCAN Clustering Algorithm
Finding clusters of different densities is a challenging task. DBSCAN “Density-Based Spatial Clustering of Applications with Noise” method has trouble discovering clusters of various densities since it uses a fixed radius. This article proposes an extended DBSCAN for finding clusters of different densities. The proposed method uses a dynamic radius and assigns a regional density value for each object, then counts the objects of similar density within the radius. If the neighborhood size ≥ MinPts, then the object is a core, and a cluster can grow from it, otherwise, the object is assigned noise temporarily. Two objects are similar in local density if their similarity ≥ threshold. The proposed method can discover clusters of any density from the data effectively. The method requires three parameters; MinPts, Eps (distance to the kth neighbor), and similarity threshold. The practical results show the superior ability of the suggested method to detect clusters of different densities even with no discernible separations between them.
Read moreEnhancing density-based clustering: Parameter reduction and outlier detection
Enhancing density-based clustering: Parameter reduction and outlier detection
An Efficient Hybrid-Clustream Algorithm for Stream Mining
Stream clustering is a standout amongst the most imperative fields in machine learning. Traditional unsupervised clustering tasks have been normally carried out in batch mode where data could be somehow fitted in memory and therefore several passes on the data are allowed. However the new Big Data paradigm has created a new environment where data can be potentially non-finite and arrive continuously. Such streams of data can reach computing systems at high speeds and contain data generation processes which might be non-stationary. For clustering tasks, this implies inconceivability to store all information in memory and obscure number and size of clusters. Noise levels can also be high due to either data generation or transmission. All these factors make traditional clustering methods not suitable to cope. As a consequence, stream clustering has emerged as a field of intense research with the aim of tackling these challenges. Clustream is one of the most advanced state of the art stream clustering algorithm. It normally requires two phases: first online micro-clustering phase, where statistics are gathered describing the incoming data; and a second offline macro-clustering phase, where a conventional non-stream clustering algorithm is executed using the high level statistics resulting from the online step. Because of its design, it requires expert-level parametrization or suffers from low runtime performance or has high sensitivity to noise or degrade considerably in high dimensional spaces because of their offline step. We propose a new stream clustering algorithm, the Clustream-hybrid based on Clustream clustering principles. It extends the same process used in Clustream but uses k-means++ instead of k-means in macro-clustering phase enabling it to accomplish quick runtime calculation while additionally keeping accuracy in high dimensional settings. We integrate it in MOA (Massive Online Analysis) tool. We evaluated the results with nine clustering quality metrics and compared the performance with Clustream for both synthetic and real data sets. The results are encproposedaging, outperforming in most of the cases in quality metrics.
Read moreMCMSTStream: applying minimum spanning tree to KD-tree-based micro-clusters to define arbitrary-shaped clusters in streaming data
Stream clustering has emerged as a vital area for processing streaming data in real-time, facilitating the extraction of meaningful information. While efficient approaches for defining and updating clusters based on similarity criteria have been proposed, outliers and noisy data within stream clustering areas pose a significant threat to the overall performance of clustering algorithms. Moreover, the limitation of existing methods in generating non-spherical clusters underscores the need for improved clustering quality. As a new methodology, we propose a new stream clustering approach, MCMSTStream, to overcome the abovementioned challenges. The algorithm applies MST to micro-clusters defined by using the KD-Tree data structure to define macro-clusters. MCMSTStream is robust against outliers and noisy data and has the ability to define clusters with arbitrary shapes. Furthermore, the proposed algorithm exhibits notable speed and can handling high-dimensional data. ARI and Purity indices are used to prove the clustering success of the MCMSTStream. The evaluation results reveal the superior performance of MCMSTStream compared to state-of-the-art stream clustering algorithms such as DenStream, DBSTREAM, and KD-AR Stream. The proposed method obtained a Purity value of 0.9780 and an ARI value of 0.7509, the highest scores for the KDD dataset. In the other 11 datasets, it obtained much higher results than its competitors. As a result, the proposed method is an effective stream clustering algorithm on datasets with outliers, high-dimensional, and arbitrary-shaped clusters. In addition, its runtime performance is also quite reasonable.
Read moreAn improved and heuristic-based iterative DBSCAN clustering algorithm
In recent years, the clustering of multi-density data has been a research hotspot. As a widely applied clustering algorithm, density-based spatial clustering of application with noise (DBSCAN) algorithm can avoid the interference of noise data, and has the ability to find clusters of any shape. However, the DBSCAN algorithm used hyperparameter (e.g., eps, MinPts), which makes the clusters of different densities cannot be identified. Furthermore, when the algorithm is applied to different datasets, the clustering quality would decrease severely. To cope with the above problems, we propose an improved and heuristic-based iterative DBSCAN clustering algorithm. The new algorithm realizes flexible clustering of the data with different densities by MinPts. In particular, we first estimate the values of MinPts according to the user area densities, then the new clustering is carried out based on the MinPts. This method avoids too dense clusters. In the experiments, Silhouette Coefficient and purity are adopted to evaluate the clustering effect of the improved DBSCAN algorithm. It is demonstrated that the proposed algorithm can effectively cluster multi-density data, and has greater adaptability and clustering performance to all kinds of data.
Read moreAn efficient and scalable density-based clustering algorithm for datasets with complex structures
An efficient and scalable density-based clustering algorithm for datasets with complex structures
On Evaluation of Data Stream Clustering Algorithms: A Survey
Data stream mining is a research area that has grown enormously in recent years. The main challenge is to extract knowledge in real-time from a possibly unbounded stream of data. Clustering, a process in which groupings within the data are to be identified, data streams is an useful technique to extract and identify underlying structures of the data. An open question in the field of stream clustering is how to evaluate the proposed algorithms. In this survey, we review the literature in the domain to identify the common methodologies, datasets, and evaluation measures, used to evaluate the algorithms. We provide a short summary of the stream clustering algorithms in the literature, but our primary focus lies in the survey of cluster validation relevant to the evaluation of data stream clustering algorithms. We begin our literature review with the inception of clustering incrementally, namely with the introduction of the balanced iterative reducing and clustering using hierarchies (BIRCH) algorithm.We identify that the evaluation methodologies primarily focus on performance, and that aspects such as cluster quality are rarely visited. Performance has been the focal point of all evaluation, both in terms of computational performance and accuracy, since the inception of clustering data streams. We also identify that issues that exist in the conventional clustering domain are also present in the data stream clustering. However, minor additions to the evaluation methods can improve both the applicability and usefulness of the algorithms.
Read moreOn Density-Based Data Streams Clustering Algorithms: A Survey
Clustering data streams has drawn lots of attention in the last few years due to their ever-growing presence. Data streams put additional challenges on clustering such as limited time and memory and one pass clustering. Furthermore, discovering clusters with arbitrary shapes is very important in data stream applications. Data streams are infinite and evolving over time, and we do not have any knowledge about the number of clusters. In a data stream environment due to various factors, some noise appears occasionally. Density-based method is a remarkable class in clustering data streams, which has the ability to discover arbitrary shape clusters and to detect noise. Furthermore, it does not need the number of clusters in advance. Due to data stream characteristics, the traditional density-based clustering is not applicable. Recently, a lot of density-based clustering algorithms are extended for data streams. The main idea in these algorithms is using density-based methods in the clustering process and at the same time overcoming the constraints, which are put out by data stream’s nature. The purpose of this paper is to shed light on some algorithms in the literature on density-based clustering over data streams. We not only summarize the main density-based clustering algorithms on data streams, discuss their uniqueness and limitations, but also explain how they address the challenges in clustering data streams. Moreover, we investigate the evaluation metrics used in validating cluster quality and measuring algorithms’ performance. It is hoped that this survey will serve as a steppingstone for researchers studying data streams clustering, particularly density-based algorithms.
Read moreAdaptive non-linear clustering in data streams
Data stream clustering has emerged as a challenging and interesting problem over the past few years. Due to the evolving nature, and one-pass restriction imposed by the data stream model, traditional clustering algorithms are inapplicable for stream clustering. This problem becomes even more challenging when the data is high-dimensional and the clusters are not linearly separable in the input space. In this paper, we propose a nonlinear stream clustering algorithm that adapts to the stream's evolutionary changes. Using the kernel methods for dealing with the non-linearity of data separation, we propose a novel 2-tier stream clustering architecture. Tier-1 captures the temporal locality in the stream, by partitioning it into segments, using a kernel-based novelty detection approach. Tier-2 exploits this segment structure to continuously project the streaming data nonlinearly onto a low-dimensional space (LDS), before assigning them to a cluster. We demonstrate the effectiveness of our approach through extensive experimental evaluation on various real-world datasets.
Read moreAn Efficient Set-Based Algorithm for Variable Streaming Clustering
In this paper, a new algorithm for Data Streaming clustering is proposed, namely the SetClust algorithm. The Data Streaming clustering model focuses on making clustering of the data while it arrives, being useful in many practical applications. The proposed algorithm, unlike other streaming clustering algorithms, is designed to handle cases when there is no available a priori information about the number of clusters to be formed, having as a second objective to discover the best number of clusters needed to represent the points. The SetClust algorithm is based on structures for disjoint-set operations, making the concept of a cluster to be the union of multiple well-formed sets to allow the algorithm to recognize non-spherical patterns even in high dimensional points. This yields to quadratic running time on the number of formed sets. The algorithm itself can be interpreted as an efficient data structure for streaming clustering. Results of the experiments show that the proposed algorithm is highly suitable for clustering quality on well-spread data points.
Read moreOptimizing Data Stream Representation: An Extensive Survey on Stream Clustering Algorithms
Analyzing data streams has received considerable attention over the past decades due to the widespread usage of sensors, social media and other streaming data sources. A core research area in this field is stream clustering which aims to recognize patterns in an unordered, infinite and evolving stream of observations. Clustering can be a crucial support in decision making, since it aims for an optimized aggregated representation of a continuous data stream over time and allows to identify patterns in large and high-dimensional data. A multitude of algorithms and approaches has been developed that are able to find and maintain clusters over time in the challenging streaming scenario. This survey explores, summarizes and categorizes a total of 51 stream clustering algorithms and identifies core research threads over the past decades. In particular, it identifies categories of algorithms based on distance thresholds, density grids and statistical models as well as algorithms for high dimensional data. Furthermore, it discusses applications scenarios, available software and how to configure stream clustering algorithms. This survey is considerably more extensive than comparable studies, more up-to-date and highlights how concepts are interrelated and have been developed over time.
Read moreAUTOCLUST+: Automatic Clustering of Point-Data Sets in the Presence of Obstacles
Wide spread clustering algorithms use the Euclidean distance to measure spatial proximity. However, obstacles in other GIS data-layers prevent traversing the straight path between two points. AUTOCLUST+ clusters points in the presence of obstacles based on Voronoi modeling and Delaunay Diagrams. The algorithm is free of usersupplied arguments and incorporates global and local variations. Thus, it detects high-quality clusters (clusters of arbitrary shapes, clusters of different densities, sparse clusters adjacent to high-density clusters, multiple bridges between clusters and closely located high-density clusters) without prior knowledge. Consequently, it successfully supports correlation analyses between layers (requiring high-quality clusters) and more general locational optimization problems in the presence of obstacles. All this within O(n log n+[m+R] log n) expected time, where n is the number of data points, m is the number of line-segments that determine the obstacles and R is the number of Delaunay edges intersecting some obstacles. A series of detailed performance evaluations illustrates the power of AUTOCLUST+ and confirms the virtues of our approach.
Read moreAn algorithm for discovering clusters of different densities or shapes in noisy data sets
In clustering spatial data, we are given a set of points in Rn and the objective is to find the clusters (representing spatial objects) in the set of points. Finding clusters with different shapes, sizes, and densities in data with noise and potentially outliers is a challenging task. This problem is especially studied in machine learning community and has lots of applications. We present a novel clustering technique, which can solve mentioned issues considerably. In the proposed algorithm, we let the structure of the data set itself find the clusters, this is done by having points actively send and receive feedbacks to each other.The idea of the proposed method is to transform the input data set into a graph by adding edges between points that belong to the same cluster, so as connected components correspond to clusters, whereas points in different clusters are almost disconnected. At the start, our algorithm creates a preliminary graph and tries to improve it iteratively. In order to build the graph (add more edges), each point sends feedback to its neighborhood points. The neighborhoods and the feedback to be sent are determined by investigating the received feedbacks. This process continues until a stable graph is created. Henceforth, the clusters are formed by post-processing the constructed graph. Our algorithm is intuitive, easy to state and analyze, and does not need to have lots of parameter tuning. Experimental results show that our proposed algorithm outperforms existing related methods in this area.
Read moreBenchmarking the benchmark — Comparing synthetic and real-world Network IDS datasets
Network Intrusion Detection Systems (NIDSs) are an increasingly important tool for the prevention and mitigation of cyber attacks. Over the past years, a lot of research efforts have aimed at leveraging the increasingly powerful models of Machine Learning (ML) for this purpose. A number of labelled synthetic datasets have been generated and made publicly available by researchers, and they have become the benchmarks via which new ML-based NIDS classifiers are being evaluated. Recently published results show excellent classification performance with these datasets, increasingly approaching 100 percent performance across key evaluation metrics such as Accuracy, F1 score, AUC, etc. Unfortunately, we have not yet seen these excellent academic research results translated into practical NIDS systems with such near-perfect performance. This motivated our research presented in this paper, where we analyse the statistical properties of the benign traffic in three of the more recent and relevant NIDS datasets, (CIC_IDS, UNSW_NB15, TON_IOT), by converting them into a common flow format. As a comparison, we consider two datasets obtained from real-world production networks, one from a university network and one from a medium size Internet Service Provider (ISP). Our results show that the two real-world datasets are quite similar among themselves in regards to most of the considered statistical features. Equally, the three synthetic datasets are also relatively similar within their group. However, and most importantly, our results show a distinct difference of most of the considered statistical features between the three synthetic datasets and the two real-world datasets. Since ML relies on the basic assumption of training and test datasets being sampled from the same distribution, this raises the question of how well the performance results of ML-classifiers trained on the considered synthetic datasets can translate and generalise to real-world networks. We believe this is an interesting and relevant question which provides motivation for further research in this space.
Read more