- Research Article
28
- 10.1016/j.knosys.2022.108266
Time-interval temporal patterns can beat and explain the malware
- Jan 29, 2022
- Knowledge-Based Systems
- Ido Finder + 2 more +2
Time-interval temporal patterns can beat and explain the malware
Malicious software also known as "Malware" is software that uses legitimate instructions or code to perform malicious actions. Malware poses a major threat for computer security and information security in general. Over the years, malware has evolved to the point that a single malware specimen can have hundreds or maybe thousands of variants using polymorphic and metamorphic transformation to change the signature of the malware variant in propagation. The common signature-based malware detection methods are no longer robust to detect these variants due to the alteration of code. Static analysis is required to obtain these signatures and anti-virus companies are required to propagate these signature updates to their software. A faster detection method is needed to compensate the exponentially increasing number of malware variants. Machine learning is a trending approach for malware detection. This removes the need to use signature-based detection and is also faster. Software companies do not need to propagate signatures as often. Machine learning algorithms using opcode sequences can recognise patterns in the malicious code that are not present in common signatures and classify them more efficiently. Therefore, a machine learning approach for malware detection should be adopted for faster and more efficient detection. Most research in malware detection using machine learning used static attributes such as network connections, processes spawned, hashes, etc., that were not that robust to changes. In this paper we introduced our novel approach in using trigrams and PE file attributes as features for malware detection. We took a text mining approach to make our detection method more robust to polymorphism and metamorphism. The instruction sequence for critical code in malware on the assembly level is basically the same across malware families. We used opcode trigram sequences as the main feature for our machine learning algorithm. We used Support Vector Machine(SVM) as our classifying algorithm which is a discriminative classifier model that gives a definite decision whether the predicted outcome belongs to the learned class or not. The above shows our novel approach that enabled us to get higher detection rates with less features.
Time-interval temporal patterns can beat and explain the malware
Time-interval temporal patterns can beat and explain the malware
A novel deep learning-based approach for malware detection
A novel deep learning-based approach for malware detection
Sequential opcode embedding-based malware detection method
Sequential opcode embedding-based malware detection method
MalGen: Malware Generation with Specific Behaviors to Improve Machine Learning-based Detectors
In recent years, infections and damage caused by malware have increased at exponential rates. At the same time, machine learning (ML) techniques have shown tremendous promise in many domains, often out performing human efforts by learning from large amounts of data. Results in the open literature suggest that ML is able to provide similar results for malware detection, achieving greater than 99% classification accuracy [49]. However, the same detection rates when applied in deployed settings have not been achieved. Malware is distinct from many other domains in which ML has shown success in that (1) it purposefully tries to hide, leading to noisy labels and (2) often its behavior is similar to benign software only differing in intent, among other complicating factors. This report details the reasons for the difficultly of detecting novel malware by ML methods and offers solutions to improve the detection of novel malware. We propose to detect malware by detecting behaviors commonly exhibited by malware such as DLL injection, and process hollowing. This is based on the assumption that there is a set of behaviors that are common to most malware samples and detecting them will generalize to novel malware. Additionally, detected behaviors point analysts toward appropriate handling and mitigation strategies, which is not the case with a binary benign/malicious classification. A behavior labeling method was developed and was used to label an existing malware dataset. Results show that detecting malicious behaviors is much more difficult than simply classifying malware and goodware?achieving 80% accuracy compared to reported 99% accuracy from classifying malware and goodware. This drop is due to several reasons which are detailed in the report. We also propose to evaluate the performance of detecting novel malware by holding out a malware family for testing and training on the other families. Traditional ML evaluation will shuffle the data and then split the data into training and testing. Our method addresses the use-case when novel malware families are encountered and they require more than just a malicious or benign designation. Our results suggest that this type of evaluation is much more difficult than traditional methods and provides more realistic results, albeit, significantly worse. For our behavior detection, accuracy decreases from 80% to 68% across all behaviors when holding out a malware family from training. We show that the degradation in performance is because each malware family has distinct characteristics resulting in high extrapolations by an ML model. Here, an ML model should return an "I do not know" response and request further analysis from an analyst. We run a number of experiments that compare novel malware families to the training data using different feature representations including a genomics-inspired distance measure and features extracted by deep learning. Generally, held-out families are significantly different from the training data, resulting in unpredictable results. This has been observed generally in the ML community [22, 9]. We empirically demonstrate this in the domain of malware detection. In an attempt to improve the detection of malware behaviors, we examine the impact that additional synthetic data has on the performance of an ML model in detecting behaviors in novel malware families. We find that while synthetic data does improve the performance of ML models, often simpler methods perform better than more complicated ones. Two generative modeling techniques were examined to produce synthetic malware samples such that the behaviors present are able to be specified externally. The difficulty is due to finer grained analysis of the executable and modifying the problem from a binary classification problem to a multi-label problem. The addition of synthetic data increases the overall accuracy from 68% to 70%. While far less accurate than measures presented in academic analyses, we believe that this is more representative of real-world performance and allows models to be properly placed within a malware detection system. We suggest that in highly dynamic environments ML pipelines should determine whether an ML model is competent in the area of new data and should involve mechanisms to improve over time with a human in the loop.
Read moreA state-of-the-art survey of malware detection approaches using data mining techniques
Data mining techniques have been concentrated for malware detection in the recent decade. The battle between security analyzers and malware scholars is everlasting as innovation grows. The proposed methodologies are not adequate while evolutionary and complex nature of malware is changing quickly and therefore turn out to be harder to recognize. This paper presents a systematic and detailed survey of the malware detection mechanisms using data mining techniques. In addition, it classifies the malware detection approaches in two main categories including signature-based methods and behavior-based detection. The main contributions of this paper are: (1) providing a summary of the current challenges related to the malware detection approaches in data mining, (2) presenting a systematic and categorized overview of the current approaches to machine learning mechanisms, (3) exploring the structure of the significant methods in the malware detection approach and (4) discussing the important factors of classification malware approaches in the data mining. The detection approaches have been compared with each other according to their importance factors. The advantages and disadvantages of them were discussed in terms of data mining models, their evaluation method and their proficiency. This survey helps researchers to have a general comprehension of the malware detection field and for specialists to do consequent examinations.
Read morePredicting and identifying factors associated with undernutrition among children under five years in Ghana using machine learning algorithms.
Undernutrition among children under the age of five is a major public health concern, especially in developing countries. This study aimed to use machine learning (ML) algorithms to predict undernutrition and identify its associated factors. Secondary data analysis of the 2017 Multiple Indicator Cluster Survey (MICS) was performed using R and Python. The main outcomes of interest were undernutrition (stunting: height-for-age (HAZ) < -2 SD; wasting: weight-for-height (WHZ) < -2 SD; and underweight: weight-for-age (WAZ) < -2 SD). Seven ML algorithms were trained and tested: linear discriminant analysis (LDA), logistic model, support vector machine (SVM), random forest (RF), least absolute shrinkage and selection operator (LASSO), ridge regression, and extreme gradient boosting (XGBoost). The ML models were evaluated using the accuracy, confusion matrix, and area under the curve (AUC) receiver operating characteristics (ROC). In total, 8564 children were included in the final analysis. The average age of the children was 926 days, and the majority were females. The weighted prevalence rates of stunting, wasting, and underweight were 17%, 7%, and 12%, respectively. The accuracies of all the ML models for wasting were (LDA: 84%; Logistic: 95%; SVM: 92%; RF: 94%; LASSO: 96%; Ridge: 84%, XGBoost: 98%), stunting (LDA: 86%; Logistic: 86%; SVM: 98%; RF: 88%; LASSO: 86%; Ridge: 86%, XGBoost: 98%), and for underweight were (LDA: 90%; Logistic: 92%; SVM: 98%; RF: 89%; LASSO: 92%; Ridge: 88%, XGBoost: 98%). The AUC values of the wasting models were (LDA: 99%; Logistic: 100%; SVM: 72%; RF: 94%; LASSO: 99%; Ridge: 59%, XGBoost: 100%), for stunting were (LDA: 89%; Logistic: 90%; SVM: 100%; RF: 92%; LASSO: 90%; Ridge: 89%, XGBoost: 100%), and for underweight were (LDA: 95%; Logistic: 96%; SVM: 100%; RF: 94%; LASSO: 96%; Ridge: 82%, XGBoost: 82%). Age, weight, length/height, sex, region of residence and ethnicity were important predictors of wasting, stunting and underweight. The XGBoost model was the best model for predicting wasting, stunting, and underweight. The findings showed that different ML algorithms could be useful for predicting undernutrition and identifying important predictors for targeted interventions among children under five years in Ghana.
Read moreMachine learning prediction in cardiovascular diseases: a meta-analysis
Several machine learning (ML) algorithms have been increasingly utilized for cardiovascular disease prediction. We aim to assess and summarize the overall predictive ability of ML algorithms in cardiovascular diseases. A comprehensive search strategy was designed and executed within the MEDLINE, Embase, and Scopus databases from database inception through March 15, 2019. The primary outcome was a composite of the predictive ability of ML algorithms of coronary artery disease, heart failure, stroke, and cardiac arrhythmias. Of 344 total studies identified, 103 cohorts, with a total of 3,377,318 individuals, met our inclusion criteria. For the prediction of coronary artery disease, boosting algorithms had a pooled area under the curve (AUC) of 0.88 (95% CI 0.84–0.91), and custom-built algorithms had a pooled AUC of 0.93 (95% CI 0.85–0.97). For the prediction of stroke, support vector machine (SVM) algorithms had a pooled AUC of 0.92 (95% CI 0.81–0.97), boosting algorithms had a pooled AUC of 0.91 (95% CI 0.81–0.96), and convolutional neural network (CNN) algorithms had a pooled AUC of 0.90 (95% CI 0.83–0.95). Although inadequate studies for each algorithm for meta-analytic methodology for both heart failure and cardiac arrhythmias because the confidence intervals overlap between different methods, showing no difference, SVM may outperform other algorithms in these areas. The predictive ability of ML algorithms in cardiovascular diseases is promising, particularly SVM and boosting algorithms. However, there is heterogeneity among ML algorithms in terms of multiple parameters. This information may assist clinicians in how to interpret data and implement optimal algorithms for their dataset.
Read moreEffects of Expression Recognition with Machine and Deep Learning Algorithms on Psychotherapy
This study presents a comparative analysis of the classification performance of facial and emotion recognition systems using Machine Learning (ML) and Deep Learning (DL) algorithms. The primary objective of this work is to evaluate the applicability of emotion recognition in fields such as psychotherapy and crime analysis, using the FER-2013 dataset.The study was conducted with ML algorithms such as Support Vector Machines (SVM), Random Forest (RF), K-Nearest Neighbor (KNN), Decision Trees, and Gradient Boosting, as well as the DL algorithm, Convolutional Neural Network (CNN). Supported by the application of feature selection, preprocessing, and feature extraction techniques, the model’s performance was measured using standard metrics such as accuracy, precision, F1-score, and AUC-ROC.The experimental results showed that the highest classification accuracy (46.61%) was achieved with the CNN model. While the ML models generally offered lower accuracy, they provided advantages in terms of computational efficiency in specific scenarios.This study aims to contribute to the literature by providing a comparative analysis of ML and DL algorithms and by highlighting the effect of data preprocessing on performance. The findings set targets for future work, such as real-time system integration and hybrid model development.
Read moreReview of Machine Learning Algorithms for Diagnosing Mental Illness
ObjectiveEnhanced technology in computer and internet has driven scale and quality of data to be improved in various areas including healthcare sectors. Machine Learning (ML) has played a pivotal role in efficiently analyzing those big data, but a general misunderstanding of ML algorithms still exists in applying them (e.g., ML techniques can settle a problem of small sample size, or deep learning is the ML algorithm). This paper reviewed the research of diagnosing mental illness using ML algorithm and suggests how ML techniques can be employed and worked in practice.MethodsResearches about mental illness diagnostic using ML techniques were carefully reviewed. Five traditional ML algorithms-Support Vector Machines (SVM), Gradient Boosting Machine (GBM), Random Forest, Naïve Bayes, and K-Nearest Neighborhood (KNN)-frequently used for mental health area researches were systematically organized and summarized.ResultsBased on literature review, it turned out that Support Vector Machines (SVM), Gradient Boosting Machine (GBM), Random Forest, Naïve Bayes, and K-Nearest Neighborhood (KNN) were frequently employed in mental health area, but many researchers did not clarify the reason for using their ML algorithm though every ML algorithm has its own advantages. In addition, there were several studies to apply ML algorithms without fully understanding the data characteristics.ConclusionResearchers using ML algorithms should be aware of the properties of their ML algorithms and the limitation of the results they obtained under restricted data conditions. This paper provides useful information of the properties and limitation of each ML algorithm in the practice of mental health.
Read moreMachine and Deep Learning Algorithms for COVID-19 Mortality Prediction Using Clinical and Radiomic Features
Aim: Machine learning (ML) and deep learning (DL) predictive models have been employed widely in clinical settings. Their potential support and aid to the clinician of providing an objective measure that can be shared among different centers enables the possibility of building more robust multicentric studies. This study aimed to propose a user-friendly and low-cost tool for COVID-19 mortality prediction using both an ML and a DL approach. Method: We enrolled 2348 patients from several hospitals in the Province of Reggio Emilia. Overall, 19 clinical features were provided by the Radiology Units of Azienda USL-IRCCS of Reggio Emilia, and 5892 radiomic features were extracted from each COVID-19 patient’s high-resolution computed tomography. We built and trained two classifiers to predict COVID-19 mortality: a machine learning algorithm, or support vector machine (SVM), and a deep learning model, or feedforward neural network (FNN). In order to evaluate the impact of the different feature sets on the final performance of the classifiers, we repeated the training session three times, first using only clinical features, then employing only radiomic features, and finally combining both information. Results: We obtained similar performances for both the machine learning and deep learning algorithms, with the best area under the receiver operating characteristic (ROC) curve, or AUC, obtained exploiting both clinical and radiomic information: 0.803 for the machine learning model and 0.864 for the deep learning model. Conclusions: Our work, performed on large and heterogeneous datasets (i.e., data from different CT scanners), confirms the results obtained in the recent literature. Such algorithms have the potential to be included in a clinical practice framework since they can not only be applied to COVID-19 mortality prediction but also to other classification problems such as diabetic prediction, asthma prediction, and cancer metastases prediction. Our study proves that the lesion’s inhomogeneity depicted by radiomic features combined with clinical information is relevant for COVID-19 mortality prediction.
Read moreA Comparative Analysis of Machine Deep Learning Algorithms for Intrusion Detection in WSN
The rapid growth of the Internet and information technologies and also the diminution of the price of hardware components like wireless sensors leads to the fast growth in Wireless Sensor Network (WSN). The WSN is a bunch/group of sensors that are located across the area, in order to monitor and determine environmental conditions such as temperature, humidity, vibrations, wind, water levels, pollution levels. WSNs are susceptible to major attacks like Denial of Service (DOS), Wormhole attack, Sinkhole attack, etc. The main threat to the WSN is because of the broadcast nature of the nodes in the network. Therefore, the security of WSNs is the essential part that must be done. Hence, to overcome these problems or threats, we are trying to detect it using AI technology. With the expanding fields of Machine Learning and Deep Learning, we can apply various algorithms in order to classify different types of attacks. Once we detect the attack properly, we can prevent it accordingly. We are using WSN-DS. It has 4 classes of attacks which are Grayhole, Blackhole, TDMA(Scheduling), and Flooding which comes under the category of DOS attacks. In this chapter, we have analysed and compared the accuracies of 5 main machine learning classification algorithms. Also, we have analysed the 1 deep learning algorithm. The ANN (Artificial Neural Network) and 5 machine learning algorithms have been trained on the dataset. Furthermore, we have used K-fold cross-validation to get even more accurate predictions. After analysis of these algorithms, we came to the point that, the Machine Learning algorithms like Random Forest, Support Vector Machines and Deep Learning algorithm, namely, Artificial Neural Network can help us for detecting intrusions in the system or network. This will help researchers in designing their own machine learning model on top of our suggested model.
Read moreOn the Evaluation of the Machine Learning Based Hybrid Approach for Android Malware Detection
Over the past few years, Android Application is deemed as one of the fastest-growing technology areas. On the other hand, the rapid growth of android applications also increases the security threats for Android users in the form of malware. Malware hacks the personal information of a user and exploits it in different criminal activities. To date, various studies have been conducted for the detection of android malware. Some authors have preferred static analysis while others have performed dynamic analysis for android malware detection. In static analysis, researchers have only employed one feature of android application either intents or permissions. As best of our knowledge, there does not exist any technique that combines both intents and permissions. In this paper, we have performed static analysis for the detection of Android malware using android intents (both implicit and explicit), android permissions and combination of intents and permissions. Furthermore, the classification is performed to analyze the effectiveness of different machine learning algorithms for malware detection and to identify the best performing feature. Our experiment results show that combination of intents and permissions play a key role in the detection of android malware.
Read moreChondrogenic Cancer Grading by Combining Machine and Deep Learning with Raman Spectra of Histopathological Tissues
Raman spectroscopy (RS) is a promising tool for cancer diagnosis. In particular, in the last years several studies have demonstrated how the diagnostic performances of RS can be significantly improved by employing machine learning (ML) algorithms for the interpretation of Raman-based data. Recently, it has been demonstrated that RS can perform an accurate classification of chondrosarcoma tissues. Chondrosarcoma is a cancer of bones, that can occur in the soft tissues near the bones. It is normally characterized by three different malignant degrees and a benign counterpart, knows as enchondroma. In line with these findings, in this paper, we exploited ML algorithms to distinguish, as well as possible, between the three grades of chondrosarcoma and to distinguish between chondrosarcoma and enchondroma. We obtained a high level of accuracy of classification by analyzing a dataset composed of a relatively small number of Raman spectra, collected in a previous study by one of the authors of this paper. Such spectra were acquired from micrometric tissue sections with a confocal Raman microscope. We tested the classification performances of a support vector machine (SVM) and a random forest classifier (RFC), as representatives of ML algorithms, and two versions of the multi-layer perceptron (MLPC) as representatives of deep learning (DL). These models, especially RFC and MLPC, showed excellent classification performances, with accuracy reaching 99.7%. This outcome makes the aforementioned models a promising route for future improvements of diagnostic devices focused on detecting cancerous bone tissues. Alongside the diagnostic purpose, the aforementioned approach allowed us to identify characteristic molecules, i.e., amino acids, nucleic acids, and bioapatites, relevant for obtaining the final diagnostic response, through the use of a tool named by us Raman Band Identification (RBI). The method to evaluate RBI is the most important contribution of this paper, because RBI could represent a relevant parameter for the identification of biochemical processes on the basis of the tumor progression and associated malignant degree. In turn, the spectral bands highlighted by RBI could provide precious indicators in an attempt to restrict the spectral acquisition to specific Raman bands. This last objective could help to reduce the amount of experimental data needed to obtain an accurate final grading outcome, with a consequent reduction in the computational cost.
Read moreIntelligent malware detection based on file relation graphs
Due to its damage to Internet security, malware and its detection has caught the attention of both anti-malware industry and researchers for decades. Many research efforts have been conducted on developing intelligent malware detection systems. In these systems, resting on the analysis of file contents extracted from the file samples, like Application Programming Interface (API) calls, instruction sequences, and binary strings, data mining methods such as Naive Bayes and Support Vector Machines have been used for malware detection. However, driven by the economic benefits, both diversity and sophistication of malware have significantly increased in recent years. Therefore, anti-malware industry calls for much more novel methods which are capable to protect the users against new threats, and more difficult to evade. In this paper, other than based on file contents extracted from the file samples, we study how file relation graphs can be used for malware detection and propose a novel Belief Propagation algorithm based on the constructed graphs to detect newly unknown malware. A comprehensive experimental study on a real and large data collection from Comodo Cloud Security Center is performed to compare various malware detection approaches. Promising experimental results demonstrate that the accuracy and efficiency of our proposed method outperform other alternate data mining based detection techniques.
Read moreSlope stability prediction using integrated metaheuristic and machine learning approaches: A comparative study
Slope stability prediction using integrated metaheuristic and machine learning approaches: A comparative study