• Home
  • Search
  • Efficient Streaming Detection of Hidden Clusters in Big Data Using Subspace Stream Clustering
  • https://doi.org/10.1007/978-3-662-43984-5_11Copy DOI Icon

Efficient Streaming Detection of Hidden Clusters in Big Data Using Subspace Stream Clustering

  • Jan 1, 2014
  • Marwan Hassani +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Abstract Recently, many data mining techniques were revisited to cope with the new big data challenges. Nearly all of these algorithms considered the efficiency of the mining algorithm to survive the increasing size of the data. However, as the dimensionality of the data increases, not only the efficiency but also the effectiveness of traditional mining algorithms is compromised. For instance, clusters hidden in some subspaces are hard to be detected using traditional clustering algorithms, as the dimensionality of the data increases. In this paper, we consider both the huge size, and the high dimensionality of big data by providing a novel solution that presents a three-phase model for subspace stream clustering algorithms. Our novel model, overcomes the huge size of the big data in its first phase, by continuously applying a streaming concept over the huge data objects, and summarizing them into micro-clusters. Then, after each certain batch of data, or after upon a user request, the second phase is applied over the data summarized in micro-clusters, to reconstruct the current distribution of the data out of the current summaries. In the third phase, a subspace clustering algorithm is applied to overcome the high dimensionality of the data, and to find hidden clusters within some subspace. An extensive evaluation study over different scenarios that follow our model over a big data set is performed.KeywordsSubspace stream clusteringStreaming big dataReal-time streaming analysis of big data

Similar Papers
  • Research Article
  • Citations4

A H-K CLUSTERING ALGORITHM FOR HIGH DIMENSIONAL DATA USING ENSEMBLE LEARNING

  • Jan 11, 2015
  • Zenodo (CERN European Organization for Nuclear Research)
  • Rashmi Paithankar
  • Research Article
  • Citations36

Fuzzy partition based soft subspace clustering and its applications in high dimensional data

  • May 28, 2013
  • Information Sciences
  • Jun Wang +3
  • Conference Article
  • Citations3

A Subtractive Based Subspace Clustering Algorithm on High Dimensional Data

  • Dec 01, 2009
  • Ying Deng +2
  • PDF
  • Research Article
  • Citations6

Research on Community Detection of Online Social Network Members Based on the Sparse Subspace Clustering Approach

  • Dec 09, 2019
  • Future Internet
  • Zihe Zhou +1
  • Conference Article

Evaluation of Big Data Privacy and Accuracy Issues

  • Jan 01, 2016
  • Reem Bashir +1
  • PDF
  • Research Article
  • Citations64

Scalable Clustering Algorithms for Big Data: A Review

  • Jan 01, 2021
  • IEEE Access
  • Mahmoud A Mahdi +2
  • Book Chapter
  • Citations4

A Data Stream Clustering Algorithm Based on Density and Extended Grid

  • Jan 01, 2017
  • Zheng Hua +3
  • Book Chapter
  • Citations2

Towards Robust Performance Guarantees for Models Learned from High-Dimensional Data

  • Jan 01, 2015
  • Rui Henriques +1
  • Book Chapter
  • Citations3

Reduct and Variance Based Clustering of High Dimensional Dataset

  • Jan 01, 2012
  • Dharmveer Singh Rajput +2
  • Book Chapter
  • Citations4

Pavement Distress Image Recognition Based on Multilayer Autoencoders

  • Jan 01, 2012
  • Lukui Shi +2
  • Research Article
  • Citations2

基于Spark的动态聚类算法研究

  • Jan 01, 2016
  • Computer Science and Application
  • 伯涛 张
  • Research Article
  • Citations27

Analyzing customer behavior from shopping path data using operation edit distance

  • Oct 13, 2016
  • Applied Intelligence
  • M Alex Syaekhoni +2
  • Research Article

Integrated framework for semantic text mining and ontology construction using inference engine

  • Jan 01, 2017
  • International Journal of Data Science
  • Purnachand Kollapudi +1
  • Research Article
  • Citations3

The Application Research of High-dimensional Mixed-attribute Data Clustering Algorithm

  • Apr 30, 2010
  • International Journal of Digital Content Technology and its Applications
  • Lilin Fan - +2
  • Research Article

Structural Optimization of Causal Driven Model Based on Bayesian Network in High-dimensional Data Classification

  • Jan 01, 2025
  • Applied Mathematics and Nonlinear Sciences
  • Kuo Li +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.