• Cite Icon6
  • https://doi.org/10.5772/6447Copy DOI Icon

Clustering Parallel Data Streams

  • Jan 1, 2009
  • Yixin Chen
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Massive volumes of data streams can be found in numerous applications such as network intrusion detection, financial transaction flows, telephone call records, sensor streams, and meteorological data. In recent years, there are increasing demands for mining data streams. Unlike the finite, statically stored data sets, stream data are massive, continuous, temporally ordered, dynamically changing, and potentially infinite [5]. For example, Cortes et al. report that AT&T long distance call records consist of 300 million records per day for 100 million customers. For the stream data applications, the volume of data is usually too huge to be stored or to be scanned for more than once. Further, in data streams, the data points can only be sequentially accessed. Random access to data is not allowed. Extensive research has been done for mining data streams, including those on the stream data classification [3, 20], mining frequent patterns [9, 17, 18], and clustering stream data [1, 2, 8, 9, 10, 11, 12, 13, 14, 16, 19]. In this paper, we study the clustering of multiple and parallel data streams. Our study should be differentiated from some previous studies on clustering stream data [19, 1]. Our goal is to group multiple streams with similar behavior and trend together, instead of to cluster the data records within one data stream. There are various applications where it is desirable to cluster the streams themselves rather than the individual data records within them. For example, the price of a stock may rise and fall from time to time. To reduce the financial risk, an investor may prefer to spread his investment over a number of stocks which may exhibit different behaviors. As another application, in meteorological study and disaster prediction, it is useful to cluster meteorological data streams from different geographical regions of similar curvature trends in order to identify regions with similar meteorological behaviors. Yet another example is that a super market may record sales on different merchandizes. There may be some relationship among the sales of different merchandizes and thus the merchant can make use of the correlation to manipulate the prices to maximize the profit. Clustering refers to partition a data set into clusters such that members within the same cluster are similar in a certain sense and members of different clusters are dissimilar. Current clustering techniques can be broadly classified into several categories: partitioning methods (e.g., k-means and k-medoids), hierarchical methods (e.g. BIRCH [22]), densitybased methods (e.g. DBSCAN [15]), and grid-based methods (e.g. CLIQUE [4]). However, these methods are designed only for static data sets and can not be directly applied to data streams. O pe n A cc es s D at ab as e w w w .in te ch w eb .o rg

Similar Papers
  • PDF
  • Research Article
  • Citations12

CC_TRS: Continuous Clustering of Trajectory Stream Data Based on Micro Cluster Life

  • Jan 01, 2017
  • Mathematical Problems in Engineering
  • Musaab Riyadh +3
  • Research Article
  • Citations202

On Density-Based Data Streams Clustering Algorithms: A Survey

  • Jan 01, 2014
  • Journal of Computer Science and Technology
  • Amineh Amini +2
  • Research Article
  • Citations63

MVStream: Multiview Data Stream Clustering.

  • Oct 29, 2019
  • IEEE Transactions on Neural Networks and Learning Systems
  • Ling Huang +3
  • Conference Article
  • Citations2

Online Embedding and Clustering of Data Streams

  • Nov 20, 2019
  • Alaettin Zubaroğlu +1
  • Book Chapter

A Dynamic Model + BFR Algorithm for Streaming Data Sorting

  • Jan 01, 2019
  • Yongwei Tan +2
  • Book Chapter
  • Citations3

Clustering Evolving Data Stream with Affinity Propagation Algorithm

  • Jan 01, 2014
  • Walid Atwa +1
  • Conference Article
  • Citations4

A Clustering Algorithm Based on Density-Grid for Stream Data

  • Dec 01, 2012
  • Dandan Zhang +6
  • PDF
  • Conference Article
  • Citations3

SLDPC: Towards Second Order Learning for Detecting Persistent Clusters in Data Streams

  • Sep 01, 2018
  • Ammar Al Abd Alazeez +2
  • Research Article
  • Citations9

A statistical approach for clustering in streaming data

  • Jan 09, 2014
  • Artificial Intelligence Research
  • Niloofar Mozafari +2
  • Conference Article
  • Citations7

A Parallel GPU-Based Approach to Clustering Very Fast Data Streams

  • Oct 17, 2015
  • Pengtao Huang +2
  • PDF
  • Research Article
  • Citations20

Learning in the presence of concept recurrence in data stream clustering

  • Sep 15, 2020
  • Journal of Big Data
  • K Namitha +1
  • Research Article
  • Citations8

Histogram-based clustering of multiple data streams

  • Mar 19, 2019
  • Knowledge and Information Systems
  • Antonio Balzanella +1
  • Book Chapter

GP Boosting Classification on Concept Drifting Data Streams

  • Jan 01, 2012
  • Dirisala J Nagendra Kumar +3
  • Research Article
  • Citations68

Mining maximal frequent patterns by considering weight conditions over data streams

  • Oct 23, 2013
  • Knowledge-Based Systems
  • Unil Yun +2
  • Book Chapter
  • Citations6

Towards a Parallel Computationally Efficient Approach to Scaling Up Data Stream Classification

  • Jan 01, 2014
  • Mark Tennant +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.