• Home
  • Search
  • Expressive and modular rule-basedclassifier for data streams
  • https://doi.org/10.48683/1926.00085883Copy DOI Icon

Expressive and modular rule-basedclassifier for data streams

Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

The advances in computing software, hardware, connected devices and wireless communication infrastructure in recent years have led to the desire to work with streaming data sources. Yet the number of techniques, approaches and algorithms which can work with data from a streaming source is still very limited, compared with batched data. Although data mining techniques have been a well-studied topic of knowledge discovery for decades, many unique properties as well as challenges in learning from a data stream have not been considered properly due to the actual presence of and the real needs to mine information from streaming data sources. This thesis aims to contribute to the knowledge by developing a rule-based algorithm to specifically learn classification rules from data streams, with the learned rules are expressive so that a human user can easily interpret the concept and rationale behind the predictions of the created model. There are two main structures to represent a classification model; the ‘tree-based’ structure and the ‘rule-based’ structure. Even though both forms of representation are popular and well-known in traditional data mining, they are different when it comes to interpretability and quality of models in certain circumstances. The first part of this thesis analyses background work and relevant topics in learning classification rules from data streams. This study provides information about the essential requirements to produce high quality classification rules from data streams and how many systems, algorithms and techniques related to learn the classification of a static dataset are not applicable in a streaming environment. The second part of the thesis investigates at a new technique to improve the efficiency and accuracy in learning heuristics from numeric features from a streaming data source. The computational cost is one of the important factors to be considered for an effective and practical learning algorithm/system because of the needs to learn from continuous arrivals of data examples sequentially and discard the seen data examples. If the computing cost is too expensive, then one may not be able to keep pace with the arrival of high velocity and possibly unbound data streams. The proposed technique was first discussed in the context of the use of Gaussian distribution as heuristics for building rule terms on numeric features. Secondly, empirical evaluation shows the successful integration of the proposed technique into an existing rule-based algorithm for the data stream, eRules. Continuing on the topic of a rule-based algorithm for classification data streams, the use of Hoeffding’s Inequality addresses another problem in learning from a data stream, namely how much data should be seen from a data stream before starting learning and how to keep the model updated over time. By incorporating the theory from Hoeffding’s Inequality, this study presents the Hoeffding Rules algorithm, which can induce modular rules directly from a streaming data source with dynamic window sizes throughout the learning period to ensure the efficiency and robustness towards the concept drifts. Concept drift is another unique challenge in mining data streams which the underlying concept of the data can change either gradually or abruptly over time and the learner should adapt to these changes as quickly as possible. This research focuses on the development of a rule-based algorithm, Hoeffding Rules, for data stream which considers streaming environments as primary data sources and addresses several unique challenges in learning rules from data streams such as concept drifts and computational efficiency. This knowledge facilitates the need and the importance of an interpretable machine learning model; applying new studies to improve the ability to mine useful insights from potentially high velocity, high volume and unbounded data streams. More broadly, this research complements the study in learning classification rules from data streams to address some of the unique challenges in data streams compared with conventional batch data, with the knowledge necessary to systematically and effectively learn expressive and modular classification rules from data streams.

Similar Papers
  • Dissertation
  • Citations8

Ontology-based access to sensor data streams

  • Jan 01, 2013
  • Jean-Paul Calbimonte
  • Research Article
  • Citations10

Detecting concept drift using HEDDM in data stream

  • Jan 01, 2019
  • International Journal of Intelligent Engineering Informatics
  • Snehlata S Dongre +2
  • Research Article
  • Citations74

Optimizing Data Stream Representation: An Extensive Survey on Stream Clustering Algorithms

  • Jan 21, 2019
  • Business & Information Systems Engineering
  • Matthias Carnein +1
  • Research Article
  • Citations24

Mining data streams with concept drifts using genetic algorithm

  • Feb 17, 2011
  • Artificial Intelligence Review
  • Periasamy Vivekanandan +1
  • Research Article

A Multistream Concept Drift Handling Framework via Data Sharing.

  • Dec 01, 2025
  • IEEE transactions on cybernetics
  • Bin Zhang +3
  • Conference Article
  • Citations9

A Noise-tolerant Fuzzy c-Means based Drift Adaptation Method for Data Stream Regression

  • Jun 01, 2019
  • Yiliao Song +3
  • Book Chapter
  • Citations5

Improving the Efficiency of Ensemble Classifier Adaptive Random Forest with Meta Level Learning for Real-Time Data Streams

  • Jan 01, 2020
  • Monika Arya +1
  • Research Article
  • Citations188

Kappa Updated Ensemble for drifting data stream mining

  • Oct 02, 2019
  • Machine Learning
  • Alberto Cano +1
  • Research Article

Technique Analysis for Multilayer Perceptrons to Deal with Concept Drift in Data Streams

  • Jan 01, 2024
  • Interdisciplinary Journal of Information, Knowledge, and Management
  • Paulo Mauricio Gonçalves Júnior +1
  • Research Article
  • Citations10

Long short-term memory self-adapting online random forests for evolving data stream regression

  • May 18, 2021
  • Neurocomputing
  • Yuan Zhong +4
  • Research Article
  • Citations429

Reacting to different types of concept drift: the Accuracy Updated Ensemble algorithm.

  • Jan 01, 2014
  • IEEE Transactions on Neural Networks and Learning Systems
  • Dariusz Brzezinski +1
  • Conference Article
  • Citations26

RDF-Gen

  • Jun 25, 2018
  • Georgios M Santipantakis +3
  • Research Article

Streaming Analytics Enhancing Predictive Traffic Incident Management

  • May 14, 2025
  • International Scientific Journal of Engineering and Management
  • Pidikiti Jahnavi
  • PDF
  • Research Article
  • Citations10

Revisiting streaming anomaly detection: benchmark and evaluation

  • Nov 07, 2024
  • Artificial Intelligence Review
  • Yang Cao +3
  • Research Article
  • Citations30

FAAD: an unsupervised fast and accurate anomaly detection method for a multi-dimensional sequence over data stream

  • Mar 01, 2019
  • Frontiers of Information Technology & Electronic Engineering
  • Bin Li +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.