A systematic literature survey of machine learning approaches to cyber data breach detection: Current research issues and future directions
Although several machine learning driven solutions are deemed to be effective at detecting data breaches, the recent proliferation in data breach incidents resulting from cyber attacks on computer networks demands an updated, thorough analysis of Machine Learning (ML) based data breach countermeasures to identify research gaps and guide future studies. In view of this, this study employs a systematic approach and draws insight from 89 research articles to classify machine learning based data breach countermeasures using eight criteria namely learning tasks, learning classifiers, datasets, feature engineering methods, multimodal approaches, pre-training approaches and performance. In classifying the studies, we: (a) propose a taxonomy of feature extraction and representation to classify studies using ten sub-criteria, (b) classify multimodal machine learning approaches used in the studies into three fusion sub-criteria: namely early fusion, intermediate fusion and late fusion, (c) show a comparison of studies based on pre-training techniques employed such as pre-text objective, learning model and data used in pre-training, (d) classify the datasets used in the study evaluation into two categories: real dataset and simulated dataset and (e) evaluate studies by detection performance and effectiveness against data breaches on unknown and obfuscated network traffic. To aid the literature identification, we analyse forty recent incidents and obtain prevalent cyber attack vectors of data breaches, which we present as the general workflow for data breaches due to cyber attacks. Finally, we highlight the research issues associated with existing ML-based data breach countermeasures and recommend future research directions.
Read more