- Research Article
143
- 10.1016/j.inffus.2021.02.009
Network traffic classification for data fusion: A survey
- Feb 12, 2021
- Information Fusion
- Jingjing Zhao + 3 more +3
Network traffic classification for data fusion: A survey
Privacy information theft traffic is usually detected using traffic classification methods, and deep learning-based detection methods are effective for this task. However, these methods have complex preprocessing processes as well as tend to ignore the deep features of network traffic, while the generalization ability is not outstanding. In this paper, a network traffic classification and detection model TCCN (Traffic Classification Capsule Network) based on capsule network is proposed. Meanwhile, PacketCGAN-based data balancing method is introduced to assist TCCN in traffic classification. A new feature graph vectorization method is used to improve the efficiency of TCCN. In addition, TCCN uses dynamic routing mechanism that can retain more valid traffic characteristics. In parallel, this paper proposes a new loss function to improve the generalization performance of TCCN. The results show that TCCN can show good performance in different experimental scenarios. After effective training, TCCN can also show high detection accuracy in new datasets, and the generalization ability of the model reaches a relatively excellent level.
Network traffic classification for data fusion: A survey
Network traffic classification for data fusion: A survey
Network Traffic Classification Model Based on Multi-scale Feature Fusion and Attention Mechanism (MSF-AttNet)
With the diversification of network traffic and the popularization of encrypted communication, traditional traffic classification methods have been difficult to cope with the challenges of new encryption protocols and hidden traffic. Deep learning, as a powerful feature extraction and classification tool, has been widely used in network traffic classification. In this paper, a deep network traffic classification model based on multi-scale feature fusion and attention mechanism (MSF-AttNet) is proposed. Features of different time scales are extracted through multi-scale convolution, and the adaptive weighting of channel attention and temporal self-attention mechanisms is combined to significantly improve the classification accuracy and interpretability of the model. Experiments show that MSF-AttNet outperforms traditional methods and other deep learning models on the ISCX VPN-nonVPN dataset.
Read moreNetwork Intrusion Detection Based on an Efficient Neural Architecture Search
Deep learning has been applied in the field of network intrusion detection and has yielded good results. In malicious network traffic classification tasks, many studies have achieved good performance with respect to the accuracy and recall rate of classification through self-designed models. In deep learning, the design of the model architecture greatly influences the results. However, the design of the network model architecture usually requires substantial professional knowledge. At present, the focus of research in the field of traffic monitoring is often directed elsewhere. Therefore, in the classification task of the network intrusion detection field, there is much room for improvement in the design and optimization of the model architecture. A neural architecture search (NAS) can automatically search the architecture of the model under the premise of a given optimization goal. For this reason, we propose a model that can perform NAS in the field of network traffic classification and search for the optimal architecture suitable for traffic detection based on the network traffic dataset. Each layer of our depth model is constructed according to the principle of maximum coding rate attenuation, which has strong consistency and symmetry in structure. Compared with some manually designed network architectures, classification indicators, such as Top-1 accuracy and F1 score, are also greatly improved while ensuring the lightweight nature of the model. In addition, we introduce a surrogate model in the search task. Compared to using the traditional NAS model to search the network traffic classification model, our NAS model greatly improves the search efficiency under the premise of ensuring that the results are not substantially different. We also manually adjust some operations in the search space of the architecture search to find a set of model operations that are more suitable for traffic classification. Finally, we apply the searched model to other traffic datasets to verify the universality of the model. Compared with several common network models in the traffic field, the searched model (NAS-Net) performs better, and the classification effect is more accurate.
Read moreAn accurate traffic classification model based on support vector machines
Network traffic classification is a fundamental research topic on high-performance network protocol design and network operation management. Compared with other state-of-the-art studies done on the network traffic classification, machine learning ML methods are more flexible and intelligent, which can automatically search for and describe useful structural patterns in a supplied traffic dataset. As a typical ML method, support vector machines SVMs based on statistical theory has high classification accuracy and stability. However, the performance of SVM classifier can be severely affected by the data scale, feature dimension, and parameters of the classifier. In this paper, a real-time accurate SVM training model named SPP-SVM is proposed. An SPP-SVM is deducted from the scaling dataset and employs principal component analysis PCA to extract data features and verify its relevant traffic features obtained from PCA. By employing PCA algorithm to do the dimension extraction, SPP-SVM confirms the critical component features, reduces the redundancy among them, and lowers the original feature dimension so as to reduce the over fitting and increase its generalization effectively. The optimal working parameters of kernel function used in SPP-SVM are derived automatically from improved particle swarm optimization algorithm, which will optimize the global solution and make its inertia weight coefficient adaptive without searching for the parameters in a wide range, traversing all the parameter points in the grid and adjusting steps gradually. The performance of its two- and multi-class classifiers is proved over 2 sets of traffic traces, coming from different topological points on the Internet. Experiments show that the SPP-SVM's two- and multi-class classifiers are superior to the typical supervised ML algorithms and performs significantly better than traditional SVM in classification accuracy, dimension, and elapsed time.
Read moreNetwork Traffic Classification Using Ensemble Learning in Software-Defined Networks
Accurate network traffic classification is essential for network management. However, existing network traffic classification methods cannot meet the demand of real networks in terms of classification performance, user privacy, latency, and control overhead. Thus, a machine learning-based approach has been used for network traffic classification. In this paper, we propose a network traffic classification framework using software-defined network (SDN) architecture. The proposed framework is entirely located in the network controller; thus, we can leverage the superior computational capacity, global visibility, and programmability of the SDN controller to realize real-time, adaptive, and accurate traffic classification. We also apply four ensemble algorithms and analyze their classification performance in terms of accuracy, precision, recall, F1-score, training time, and classification time. The experimental results reveal that ensemble model-based network traffic classifiers outperform other classifiers based on the proposed framework and the real-world network traffic dataset. Notably, the LightGBM model achieves the best classification performance.
Read moreA Comparative Study of Traffic Classification Techniques for Smart City Networks
Smart city networks involve many applications that impose specific Quality of Service (QoS) requirements, thus representing a challenging scenario for network management. Solutions aiming to guarantee QoS support have not been deployed in large-scale networks. Traffic classification is a mechanism used to manage different aspects, including QoS requirements. However, conventional traffic classification methods, such as the port-based method, are inefficient because of their inability to handle dynamic port allocation and encryption. Traffic classification using machine learning has gained research interest as an alternative method to achieve high performance. In fact, machine learning embeds intelligence into network functions, thus improving network management. In this study, we apply machine learning algorithms to predict network traffic classification. We apply four supervised learning algorithms: support vector machine, random forest, k-nearest neighbors, and decision tree. We also apply a port-based method of traffic classification based on applications’ popular assigned port numbers. Then, we compare the results of this method to those obtained from the machine learning algorithms. The evaluation results indicate that the decision tree algorithm provides the highest average accuracy among the evaluated algorithms, at 99.18%. Moreover, network traffic classification using machine learning provides more accurate results and higher performance than the port-based method.
Read moreNot Afraid of the Unseen: a Siamese Network based Scheme for Unknown Traffic Discovery
As an essential task for network management and security, network traffic classification has attracted increasing attention in recent years. Traditional traffic classification methods achieve certain success in identifying specific application traffic but fail with un-predefined unknown classes. Existing unknown traffic discovery methods commonly pick out some unlabeled testing data as part of training data to train the classification models, which is not in line with the real-world open environments. In this paper, we propose a novel scheme named SEEN to achieve unknown traffic detection in network traffic classification. There are three crucial phases in the SEEN: unknown discovery, unknown clustering, and system update. In the first step, using a metric-based approach with siamese network, SEEN identifies unknown traffic as well as accurately classifies the traffic generated by pre-defined application classes. After discovery, unknown traffic is automatically clustered into more fine-grained categories in the unknown clustering step. In the system update step, inspired by low-shot learning, SEEN allows new classes to be added or unnecessary known classes to be deleted quickly without retraining from the sketch, which can complement the system’s knowledge. Experimental results exhibit that SEEN can achieve outstanding performances both on known and unknown traffic identification on two open real-world datasets, and the proposed scheme can address the problem of unknown traffic effectively.
Read moreImproved EM method for internet traffic classification
Network traffic classification algorithm based on the machine learning has attracted more and more attention. Because the traditional EM algorithm has the disadvantage that the algorithm has the sensitivity of initial value and converge to local optimal point easily. This paper proposed a new improved EM algorithm based on the q-DAEM. The improved algorithm applies the EM algorithm to generate a constrained matrix, then combine the constrained matrix with the q-DAEM algorithm to reduce the search range, so that a better Gaussian mixture model can be derived from this algorithm. The algorithm was applied to the Moore datasets for evaluation, the experimental results show that this improved algorithm which applied to network traffic classification can lead to a higher precision and overall accuracy.
Read moreFLITC: A Novel Federated Learning-Based Method for IoT Traffic Classification
Internet of Things (IoT) systems are rightly receiving considerable interest for many real-world applications, from in-body networks to satellite networks. Such a massive-scale system generates a considerable amount of traffic data, making IoT systems a distributed data source generator. For many reasons, such as the functionality of IoT applications and Quality of Service (QoS) provisioning, classifying these traffic data is of high importance. In the last few years, widespread interest has been expressed in applying Machine Learning (ML)-based techniques for Network Traffic Classification (NTC) tasks. However, the traditional centralized learning-based traffic classifiers pose serious challenges, especially in IoT networks. The centralized ML techniques call for collecting a large amount of data from various IoT devices, which in turn introduces data governance and privacy challenges. Furthermore, in the centralized ML, training data need to be transferred to the Cloud, which increases communication cost and latency. To address these problems, we propose Federated Learning (FL) Internet of Things (IoT) Traffic Classifier (FLITC)-a Federated Learning (FL)-based IoT traffic classification method which is based on the Multi-Layer Perception (MLP) neural network and holds the local data unimpaired on IoT devices by sending only the learned parameters to the aggregation server. Our experimental results show that the FLITC beats centralized learning in preserving the privacy of sensitive data and offers a better degree of accuracy at the cost of a longer training time.
Read moreByteSGAN: A semi-supervised Generative Adversarial Network for encrypted traffic classification in SDN Edge Gateway
ByteSGAN: A semi-supervised Generative Adversarial Network for encrypted traffic classification in SDN Edge Gateway
Radio Frequency Traffic Classification Over WLAN
Network traffic classification is the process of analyzing traffic flows and associating them to different categories of network applications. Network traffic classification represents an essential task in the whole chain of network security. Some of the most important and widely spread applications of traffic classification are the ability to classify encrypted traffic, the identification of malicious traffic flows, and the enforcement of security policies on the use of different applications. Passively monitoring a network utilizing low-cost and low-complexity wireless local area network (WLAN) devices is desirable. Mobile devices can be used or existing office desktops can be temporarily utilized when their computational load is low. This reduces the burden on existing network hardware. The aim of this paper is to investigate traffic classification techniques for wireless communications. To aid with intrusion detection, the key goal is to passively monitor and classify different traffic types over WLAN to ensure that network security policies are adhered to. The classification of encrypted WLAN data poses some unique challenges not normally encountered in wired traffic. WLAN traffic is analyzed for features that are then used as an input to six different machine learning (ML) algorithms for traffic classification. One of these algorithms (a Gaussian mixture model incorporating a universal background model) has not been applied to wired or wireless network classification before. The authors also propose a ML algorithm that makes use of the well-known vector quantization algorithm in conjunction with a decision tree—referred to as a TRee Adaptive Parallel Vector Quantiser. This algorithm has a number of advantages over the other ML algorithms tested and is suited to wireless traffic classification. An average F-score (harmonic mean of precision and recall) > 0.84 was achieved when training and testing on the same day across six distinct traffic types.
Read morePayload-Based Traffic Classification Using Multi-Layer LSTM in Software Defined Networks
Recently, with the advent of various Internet of Things (IoT) applications, a massive amount of network traffic is being generated. A network operator must provide different quality of service, according to the service provided by each application. Toward this end, many studies have investigated how to classify various types of application network traffic accurately. Especially, since many applications use temporary or dynamic IP or Port numbers in the IoT environment, only payload-based network traffic classification technology is more suitable than the classification using the packet header information as well as payload. Furthermore, to automatically respond to various applications, it is necessary to classify traffic using deep learning without the network operator intervention. In this study, we propose a traffic classification scheme using a deep learning model in software defined networks. We generate flow-based payload datasets through our own network traffic pre-processing, and train two deep learning models: 1) the multi-layer long short-term memory (LSTM) model and 2) the combination of convolutional neural network and single-layer LSTM models, to perform network traffic classification. We also execute a model tuning procedure to find the optimal hyper-parameters of the two deep learning models. Lastly, we analyze the network traffic classification performance on the basis of the F1-score for the two deep learning models, and show the superiority of the multi-layer LSTM model for network packet classification.
Read moreEvaluation of feature selection on network traffic classification
Malicious traffic classification has become a challenge in modern communications. It is a very important task for a trained model to successfully distinguish malicious traffic. With the gradual application of machine learning and deep learning in the field of traffic classification, traffic classification has reached a high accuracy rate. Feature selection can lighten models and improve classification performance by selecting the optimal sub-feature set. Therefore, the selection of effective features is an important issue for malicious traffic classification. In this article, we propose the idea of applying feature selection methods Information Gain and RFE to malicious traffic classification. The essence is to select an effective and optimal sub-feature set from a large number of features to characterize network traffic. Then, we used the deep learning method CNN and the machine learning method RF on the three real network traffic datasets of CICIDS2017, NSL-KDD and UNSW-NB15 respectively to evaluate and verify. The experiment shows that the combination of CNN and Information Gain has the best effect. The results of many experiments show that the performance of traffic classification is greatly improved after feature selection.
Read moreAn Improved Capsule Network for Image Classification Using Multi-Scale Feature Extraction
In the realm of image classification, the capsule network is a network topology that packs the extracted features into many capsules, performs sophisticated capsule screening using a dynamic routing mechanism, and finally recognizes that each capsule corresponds to a category feature. Compared with previous network topologies, the capsule network has more sophisticated operations, uses a large number of parameter matrices and vectors to express picture attributes, and has more powerful image classification capabilities. However, in the practical application field, the capsule network has always been constrained by the quantity of calculation produced by the complicated structure. In the face of basic datasets, it is prone to over-fitting and poor generalization and often cannot satisfy the high computational overhead when facing complex datasets. Based on the aforesaid problems, this research proposes a novel enhanced capsule network topology. The upgraded network boosts the feature extraction ability of the network by incorporating a multi-scale feature extraction module based on proprietary star structure convolution into the standard capsule network. At the same time, additional structural portions of the capsule network are changed, and a variety of optimization approaches such as dense connection, attention mechanism, and low-rank matrix operation are combined. Image classification studies are carried out on different datasets, and the novel structure suggested in this paper has good classification performance on CIFAR-10, CIFAR-100, and CUB datasets. At the same time, we also achieved 98.21% and 95.38% classification accuracy on two complicated datasets of skin cancer ISIC derived and Forged Face EXP.
Read moreExtraction of Minimal Set of Traffic Features Using Ensemble of Classifiers and Rank Aggregation for Network Intrusion Detection Systems
Network traffic classification models, an essential part of intrusion detection systems, need to be as simple as possible due to the high speed of network transmission. One of the fastest approaches is based on decision trees, where the classification process requires a series of tests, resulting in a class assignment. In the network traffic classification process, these tests are performed on extracted traffic features. The classification computational efficiency grows when the number of features and their tests in the decision tree decreases. This paper investigates the relationship between the number of features used to construct the decision-tree-based intrusion detection model and the classification quality. This work deals with a reference dataset that includes IoT/IIoT network traffic. A feature selection process based on the aggregated rank of features computed as the weighted average of rankings obtained using multiple (in this case, six) classifier-based feature selectors is proposed. It results in a ranking of 32 features sorted by importance and usefulness in the classification process. In the outcome of this part of the study, it turns out that acceptable classification results for the smallest number of best features are achieved for the eight most important features at −95.3% accuracy. In the second part of these experiments, the dependence of the classification speed and accuracy on the number of most important features taken from this ranking is analyzed. In this investigation, optimal times are also obtained for eight or fewer number of the most important features, e.g., the trained decision tree needs 0.95 s to classify nearly 7.6 million samples containing eight network traffic features. The conducted experiments prove that a subset of just a few carefully selected features is sufficient to obtain reasonably high classification accuracy and computational efficiency.
Read more