- Research Article
17
- 10.1016/j.jocs.2022.101925
A varied density-based clustering algorithm
- Dec 12, 2022
- Journal of Computational Science
- Ahmed Fahim
A varied density-based clustering algorithm
Several clustering algorithms have been extensively used to analyze vast amounts of spatial data. One of these algorithms is the SNN (Shared Nearest Neighbor), a density-based algorithm, which has several advantages when analyzing this type of data due to its ability of identifying clusters of different shapes, sizes and densities, as well as the capability to deal with noise. Having into account that data are usually progressively collected as time passes, incremental clustering approaches are required when there is the need to update the clustering results as new data become available. This paper proposes SNN <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">++</sup> , an incremental clustering algorithm based on the SNN. Its performance and the quality of the resulting clusters are compared with the SNN and the results show that the SNN <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">++</sup> yields the same result as the SNN and show that the incremental feature was added to the SNN without any computational penalty. Moreover, the experimental results also show that processing huge amounts of data using increments considerably decreases the number of distances that need to be computed to identify the points' nearest neighbors.
A varied density-based clustering algorithm
A varied density-based clustering algorithm
An Exhaustive Research on the Application of Intrusion Detection Technology in Computer Network Security in Sensor Networks
Intrusion detection is crucial in computer network security issues; therefore, this work is aimed at maximizing network security protection and its improvement by proposing various preventive techniques. Outlier detection and semisupervised clustering algorithms based on shared nearest neighbors are proposed in this work to address intrusion detection by converting it into a problem of mining outliers using the network behavior dataset. The algorithm uses shared nearest neighbors as similarity, judges whether it is an outlier according to the number of nearest neighbors of a data point, and performs semisupervised clustering on the dataset where outliers are deleted. In the process of semisupervised clustering, vast prior knowledge is added, and the dataset is clustered according to the principle of graph segmentation. The novelty of the proposed algorithm lies in outlier detection while effectively avoiding the dependence on parameters, thus eliminating the influence of outliers on clustering. This article uses real datasets: lypmphography and glass for simulation purposes. The simulation results show that the algorithm proposed in this paper can effectively detect outliers and has a good clustering effect. Furthermore, the experimentation reveals that the outlier detection‐based SCA‐SNN algorithm has the best practical effect on the dataset without outliers, clearly validating the clustering performance of the outlier detection‐based SCA‐SNN algorithm. Furthermore, compared to the other state‐of‐the‐art anomaly detection method, it was revealed that the anomaly detection technology based on outlier mining does not require a training process. Thus, they overcome the current anomaly detection problems caused due to incomplete normal patterns in training samples.
Read moreAn Improved Random Seed Searching Clustering Algorithm Based on Shared Nearest Neighbor
Clustering analysis continually consider as a hot field in Data Mining. For different types data sets and application purposes, the relevant researchers concern on various aspect, such as the adaptability to fit density and shape, noise detection, outliers identification, cluster number determination, accuracy and optimization. Lots of related works focus on the Shared Nearest Neighbor measure method, due to its best and wide adaptability to deal with complex distribution data set. Based on Shared Nearest Neighbor, an improved algorithm is proposed in this paper, it mainly target on the problems solution of natural distribute density, arbitrary shape and cluster number determination. The new algorithm start with random selected seed, follow the direction of its nearest neighbors, search and find its neighbors which have the greatest similar features, form the local maximum cluster, dynamically adjust the data objects’ affiliation to realize the local optimization at the same time, and then end the clustering procedure until identify all the data objects. Experiments verify the new algorithm has the advanced ability to fit the problems such as different density, shape, noise, cluster number and so on, and can realize fast optimization searching.
Read moreDiscovery of patterns in spatio-temporal data using clustering techniques
Spatial-temporal clustering is very useful unsupervised learning technique and can be used to identify interesting distribution patterns from geo-located data. It is one of the most commonly used data mining techniques in many application domains, e.g. geographic information science, health science, and environmental science. In this paper, we propose a density-based spatial-temporal clustering algorithm for geo-located data points, based on an extension of the SNN (Shared Nearest Neighbor) clustering. The proposed algorithm allows the integration of location, time and other semantic attributes in the clustering process. This algorithm can find clusters of different sizes, shapes, and densities in noisy data. We evaluate the effectiveness of our algorithm through a case study involving a New York City taxi cab pickup data and Maryland crime data. The experimental results show that the proposed algorithm can discover interesting patterns and useful information from spatial-temporal data.
Read moreOutlier Robust Geodesic K-means Algorithm for High Dimensional Data
This paper proposes an outlier robust geodesic K-mean algorithm for high dimensional data. The proposed algorithm features three novel contributions. First, it employs a shared nearest neighbour (SNN) based distance metric to construct the nearest neighbour data model. Second, it combines the notion of geodesic distance to the well-known local outlier factor (LOF) model to distinguish outliers from inlier data. Third, it introduces a new ad-hoc strategy to integrate outlier scores into geodesic distances. Numerical experiments with synthetic and real world remote sensing spectral data show the efficiency of the proposed algorithm in clustering of high-dimensional data in terms of the overall clustering accuracy and the average precision.
Read moreShared Nearest Neighbor Clustering in a Locality Sensitive Hashing Framework.
We present a new algorithm to cluster high-dimensional sequence data and its application to the field of metagenomics, which aims at reconstructing individual genomes from a mixture of genomes sampled from an environmental site, without any prior knowledge of reference data (genomes) or the shape of clusters. Such problems typically cannot be solved directly with classical approaches seeking to estimate the density of clusters, for example, using the shared nearest neighbors (SNN) rule, due to the prohibitive size of contemporary sequence datasets. We explore here a new approach based on combining the SNN rule with the concept of locality sensitive hashing (LSH). The proposed method, called LSH-SNN, works by randomly splitting the input data into smaller-sized subsets (buckets) and employing the SNN rule on each of these buckets. Links can be created among neighbors sharing a sufficient number of elements, hence allowing clusters to be grown from linked elements. LSH-SNN can scale up to larger datasets consisting of millions of sequences, while achieving high accuracy across a variety of sample sizes and complexities.
Read moreAn Unsupervised Hyperspectral Band Selection Method Based on Shared Nearest Neighbor and Correlation Analysis
Band selection is an important dimensionality reduction (DR) methodology for hyperspectral images (HSI). In recent years, many ranking-based clustering band selection methods have been developed. However, these methods do not consider the combination of bands in different clusters but only select the desired number of clustering centers based on band ranking to construct the reduced band subset, which may lead to obtaining a set of bands with low redundancy but little information or a set of bands with a large amount of information but high redundancy, thus falling into the local optimal solution set. To solve this problem, an unsupervised hyperspectral band selection method based on shared nearest neighbor and correlation analysis (SNNCA) is proposed in this paper. The proposed SNNCA method considers the interaction of bands in different clusters, and can obtain a set of bands with a large amount of information and low redundancy. First, this method uses the shared nearest neighbor to describe the local density of each band and takes the product of local density and distance factor as the weight to rank each band to select the required number of clustering centers, which ensures low redundancy among the clustering centers. Then, all bands are grouped into several clusters based on the Euclidean distance matrix and the clustering centers. Finally, the correlation among intra-cluster and inter-cluster bands and the information entropy are further analyzed, and the most representative band is selected from each cluster. The experimental results on two HSI datasets demonstrate that the proposed SNNCA method achieves better classification performance than that of other state-of-the-art comparison methods and possesses competitive running time.
Read moreA Fault Detection and Isolation Method via Shared Nearest Neighbor for Circulating Fluidized Bed Boiler
Accurate and timely fault detection and isolation (FDI) improve the availability, safety, and reliability of target systems and enable cost-effective operations. In this study, a shared nearest neighbor (SNN)-based method is proposed to identify the fault variables of a circulating fluidized bed boiler. SNN is a derivative method of the k-nearest neighbor (kNN), which utilizes shared neighbor information. The distance information between these neighbors can be applied to FDI. In particular, the proposed method can effectively detect faults by weighing the distance values based on the number of neighbors they share, thereby readjusting the distance values based on the shared neighbors. Moreover, the data distribution is not constrained; therefore, it can be applied to various processes. Unlike principal component analysis and independent component analysis, which are widely used to identify fault variables, the main advantage of SNN is that it does not suffer from smearing effects, because it calculates the contributions from the original input space. The proposed method is applied to two case studies and to the failure case of a real circulating fluidized bed boiler to confirm its effectiveness. The results show that the proposed method can detect faults earlier (1 h 39 min 46 s) and identify fault variables more effectively than conventional methods.
Read moreDTI-SNNFRA: Drug-target interaction prediction by shared nearest neighbors and fuzzy-rough approximation
In-silico prediction of repurposable drugs is an effective drug discovery strategy that supplements de-nevo drug discovery from scratch. Reduced development time, less cost and absence of severe side effects are significant advantages of using drug repositioning. Most recent and most advanced artificial intelligence (AI) approaches have boosted drug repurposing in terms of throughput and accuracy enormously. However, with the growing number of drugs, targets and their massive interactions produce imbalanced data which may not be suitable as input to the classification model directly. Here, we have proposed DTI-SNNFRA, a framework for predicting drug-target interaction (DTI), based on shared nearest neighbour (SNN) and fuzzy-rough approximation (FRA). It uses sampling techniques to collectively reduce the vast search space covering the available drugs, targets and millions of interactions between them. DTI-SNNFRA operates in two stages: first, it uses SNN followed by a partitioning clustering for sampling the search space. Next, it computes the degree of fuzzy-rough approximations and proper degree threshold selection for the negative samples’ undersampling from all possible interaction pairs between drugs and targets obtained in the first stage. Finally, classification is performed using the positive and selected negative samples. We have evaluated the efficacy of DTI-SNNFRA using AUC (Area under ROC Curve), Geometric Mean, and F1 Score. The model performs exceptionally well with a high prediction score of 0.95 for ROC-AUC. The predicted drug-target interactions are validated through an existing drug-target database (Connectivity Map (Cmap)).
Read moreHeuristic Planning Method of EV Fast Charging Station on a Freeway Considering the Power Flow Constraints of the Distribution Network
Heuristic Planning Method of EV Fast Charging Station on a Freeway Considering the Power Flow Constraints of the Distribution Network
Read moreSeasonal prediction of summer monsoon rainfall over cluster regions of India
Shared nearest neighbour (SNN) cluster algorithm has been applied to seasonal (June–September) rainfall departures over 30 sub-divisions of India to identify the contiguous homogeneous cluster regions over India. Five cluster regions are identified. Rainfall departure series for these cluster regions are prepared by area weighted average rainfall departures over respective sub-divisions in each cluster. The interannual and decadal variability in rainfall departures over five cluster regions is discussed. In order to consider the combined effect of North Atlantic Oscillation (NAO) and Southern Oscillation (SO), an index called effective strength index (ESI) has been defined. It has been observed that the circulation is drastically different in positive and negative phases of ESI-tendency from January to April. Hence, for each phase of ESI-tendency (positive and negative), separate prediction models have been developed for predicting summer monsoon rainfall over identified clusters. The performance of these models have been tested and found to be encouraging.
Read moreIncremental Spectral Clustering With Application to Monitoring of Evolving Blog Communities
Previous chapter Next chapter Full AccessProceedings Proceedings of the 2007 SIAM International Conference on Data Mining (SDM)Incremental Spectral Clustering With Application to Monitoring of Evolving Blog CommunitiesHuazhong Ning, Wei Xu, Yun Chi, Yihong Gong, and Thomas HuangHuazhong Ning, Wei Xu, Yun Chi, Yihong Gong, and Thomas Huangpp.261 - 272Chapter DOI:https://doi.org/10.1137/1.9781611972771.24PDFBibTexSections ToolsAdd to favoritesExport CitationTrack CitationsEmail SectionsAboutAbstract In recent years, spectral clustering method has gained attentions because of its superior performance compared to other traditional clustering algorithms such as K-means algorithm. The existing spectral clustering algorithms are all off-line algorithms, i.e., they can not incrementally update the clustering result given a small change of the data set. However, the capability of incrementally updating is essential to some applications such as real time monitoring of the evolving communities of websphere or blogsphere. Unlike traditional stream data, these applications require incremental algorithms to handle not only insertion/deletion of data points but also similarity changes between existing items. This paper extends the standard spectral clustering to such evolving data by introducing the incidence vector/matrix to represent two kinds of dynamics in the same framework and by incrementally updating the eigenvalue system. Our incremental algorithm, initialized by a standard spectral clustering, continuously and efficiently updates the eigenvalue system and generates instant cluster labels, as the data set is evolving. The algorithm is applied to a blog data set. Compared with recomputation of the solution by standard spectral clustering, it achieves similar accuracy but with much lower computational cost. Close inspection into the blog content shows that the incremental approach can discover not only the stable blog communities but also the evolution of the individual multi-topic blogs. Previous chapter Next chapter RelatedDetails Published:2007ISBN:978-0-89871-630-6eISBN:978-1-61197-277-1 https://doi.org/10.1137/1.9781611972771Book Series Name:ProceedingsBook Code:PR127Book Pages:xiv + 648Key words:Incremental clustering, Spectral Clustering, Incidence Vector/Matrix, Web-blogs
Read moreSingle cell clustering based on cell-pair differentiability correlation and variance analysis.
The rapid advancement of single cell technologies has shed new light on the complex mechanisms of cellular heterogeneity. Identification of intercellular transcriptomic heterogeneity is one of the most critical tasks in single-cell RNA-sequencing studies. We propose a new cell similarity measure based on cell-pair differentiability correlation, which is derived from gene differential pattern among all cell pairs. Through plugging into the framework of hierarchical clustering with this new measure, we further develop a variance analysis based clustering algorithm 'Corr' that can determine cluster number automatically and identify cell types accurately. The robustness and superiority of the proposed algorithm are compared with representative algorithms: shared nearest neighbor (SNN)-Cliq and several other state-of-the-art clustering methods, on many benchmark or real single cell RNA-sequencing datasets in terms of both internal criteria (clustering number and accuracy) and external criteria (purity, adjusted rand index, F1-measure). Moreover, differentiability vector with our new measure provides a new means in identifying potential biomarkers from cancer related single cell datasets even with strong noise. Prognosis analyses from independent datasets of cancers confirmed the effectiveness of our 'Corr' method. The source code (Matlab) is available at http://sysbio.sibcb.ac.cn/cb/chenlab/soft/Corr--SourceCodes.zip. Supplementary data are available at Bioinformatics online.
Read moreRobust Feature-Preserving Mesh Denoising Based on Consistent Subneighborhoods
In this paper, we introduce a feature-preserving denoising algorithm. It is built on the premise that the underlying surface of a noisy mesh is piecewise smooth, and a sharp feature lies on the intersection of multiple smooth surface regions. A vertex close to a sharp feature is likely to have a neighborhood that includes distinct smooth segments. By defining the consistent subneighborhood as the segment whose geometry and normal orientation most consistent with those of the vertex, we can completely remove the influence from neighbors lying on other segments during denoising. Our method identifies piecewise smooth subneighborhoods using a robust density-based clustering algorithm based on shared nearest neighbors. In our method, we obtain an initial estimate of vertex normals and curvature tensors by robustly fitting a local quadric model. An anisotropic filter based on optimal estimation theory is further applied to smooth the normal field and the curvature tensor field. This is followed by second-order bilateral filtering, which better preserves curvature details and alleviates volume shrinkage during denoising. The support of these filters is defined by the consistent subneighborhood of a vertex. We have applied this algorithm to both generic and CAD models, and sharp features, such as edges and corners, are very well preserved.
Read moreVDMR-DBSCAN: Varied Density MapReduce DBSCAN
DBSCAN is a well-known density based clustering algorithm, which can discover clusters of different shapes and sizes along with outliers. However, it suffers from major drawbacks like high computational cost, inability to find varied density clusters and dependency on user provided input density parameters. To address these issues, we propose a novel density based clustering algorithm titled, VDMR-DBSCAN (Varied Density MapReduce DBSCAN), a scalable DBSCAN algorithm using MapReduce which can detect varied density clusters with automatic computation of input density parameters. VDMR-DBSCAN divides the data into small partitions which are parallely processed on Hadoop platform. Thereafter, density variations in a partition are analyzed statistically to divide the data into groups of similar density called Density level sets (DLS). Input density parameters are estimated for each DLS, later DBSCAN is applied on each DLS using its corresponding density parameters. Most importantly, we propose a novel merging technique, which merges the similar density clusters present in different partitions and produces meaningful and compact clusters of varied density. We experimented on large and small synthetic datasets which well confirms the efficacy of our algorithm in terms of scalability and ability to find varied density clusters.
Read more