- Research Article
280
- 10.1016/j.neunet.2024.106207
A Comprehensive Survey on Deep Graph Representation Learning
- Feb 27, 2024
- Neural Networks
- Wei Ju + 15 more +15
A Comprehensive Survey on Deep Graph Representation Learning
Increasing numbers of software vulnerabilities are discovered every year whether they are reported publicly or discovered internally in proprietary code. These vulnerabilities can pose serious risk of exploit and result in system compromise, information leaks, or denial of service. We leveraged the wealth of C and C++ open-source code available to develop a largescale function-level vulnerability detection system using machine learning. To supplement existing labeled vulnerability datasets, we compiled a vast dataset of millions of open-source functions and labeled it with carefully-selected findings from three different static analyzers that indicate potential exploits. Using these datasets, we developed a fast and scalable vulnerability detection tool based on deep feature representation learning that directly interprets lexed source code. We evaluated our tool on code from both real software packages and the NIST SATE IV benchmark dataset. Our results demonstrate that deep feature representation learning on source code is a promising approach for automated software vulnerability detection.
A Comprehensive Survey on Deep Graph Representation Learning
A Comprehensive Survey on Deep Graph Representation Learning
Image cyberbullying detection and recognition using transfer deep machine learning
Image cyberbullying detection and recognition using transfer deep machine learning
Android Malware Detection Using Supervised Deep Graph Representation Learning
Despite the continuous evolution and significant improvement of cybersecurity mechanisms, malware threats remain one of the most important concerns in cyberspace. Meanwhile, Android malware plays a big role in these ever-growing threats. In recent years, deep learning has become the dominant machine learning technique for malware detection and continues to make outstanding achievements. Deep graph representation learning is the task of embedding graph-structured data into a low-dimensional space using deep learning models. Recently, autoencoders have proven to be an effective way for deep representation learning. However, it is not straightforward to apply the idea of autoencoder to graph-structured data because of their irregular structure. In this paper, we present DroidMalGNN, a novel deep learning technique that combines autoencoders with graph neural networks (GNNs) to detect Android malware in an end-to-end manner. DroidMalGNN represents each Android application with an attributed function call graph (AFCG) that allows it to model complex relationships between data. For more efficiency, DroidMalGNN performs graph representation learning in a supervised manner where two autoencoders are trained with benign and malicious AFCGs separately. In this way, it generates two informative embedding vectors for each AFCG in a low-dimensional space and feeds them into a dense neural network to classify the AFCG as benign or malicious. Our experimental results show that DroidMalGNN can achieve good detection performance in terms of different evaluation measures.
Read moreCentroids-guided deep multi-view K-means clustering
Centroids-guided deep multi-view K-means clustering
A deep multi-task representation learning method for time series classification and retrieval
A deep multi-task representation learning method for time series classification and retrieval
Perceptual-Similarity-Aware Deep Speaker Representation Learning for Multi-Speaker Generative Modeling
We propose novel deep speaker representation learning that considers perceptual similarity among speakers for multi-speaker generative modeling. Following its success in accurate discriminative modeling of speaker individuality, knowledge of deep speaker representation learning (i.e., speaker representation learning using deep neural networks) has been introduced to multi-speaker generative modeling. However, the conventional discriminative algorithm does not necessarily learn speaker embeddings suitable for such generative modeling, which may result in lower quality and less controllability of synthetic speech. We propose three representation learning algorithms that utilize a perceptual speaker similarity matrix obtained by large-scale perceptual scoring of speaker-pair similarity. The algorithms train a speaker encoder to learn speaker embeddings with three different representations of the matrix: a set of vectors, the Gram matrix, and a graph. Furthermore, we propose an active learning algorithm that iterates the perceptual scoring and speaker encoder training. To obtain accurate embeddings while reducing costs of scoring and training, the algorithm selects unscored speaker-pairs to be scored next on the basis of the sequentially-trained speaker encoder's similarity prediction results. Experimental evaluation results show that 1) the proposed representation learning algorithms learn speaker embeddings strongly correlated with perceptual speaker-pair similarity, 2) the embeddings improve synthetic speech quality in speech autoencoding tasks better than conventional d-vectors learned by discriminative modeling, 3) the proposed active learning algorithm achieves higher synthetic speech quality while reducing costs of scoring and training, and 4) among the proposed similarity {vector, matrix, graph} embedding algorithms, the first achieves the best speaker similarity for synthetic speech and the third gives the most improvement in the synthetic speech naturalness.
Read moreTri-party deep network representation learning using inductive matrix completion
Most existing network representation learning algorithms focus on network structures for learning. However, network structure is only one kind of view and feature for various networks, and it cannot fully reflect all characteristics of networks. In fact, network vertices usually contain rich text information, which can be well utilized to learn text-enhanced network representations. Meanwhile, Matrix-Forest Index (MFI) has shown its high effectiveness and stability in link prediction tasks compared with other algorithms of link prediction. Both MFI and Inductive Matrix Completion (IMC) are not well applied with algorithmic frameworks of typical representation learning methods. Therefore, we proposed a novel semi-supervised algorithm, tri-party deep network representation learning using inductive matrix completion (TDNR). Based on inductive matrix completion algorithm, TDNR incorporates text features, the link certainty degrees of existing edges and the future link probabilities of non-existing edges into network representations. The experimental results demonstrated that TFNR outperforms other baselines on three real-world datasets. The visualizations of TDNR show that proposed algorithm is more discriminative than other unsupervised approaches.
Read moreMulti-spectral Palmprint Recognition with Deep Multi-view Representation Learning
With the widespread application of biometrics in identification systems, palmprint recognition technology, as an emerging biometric technology, has received more and more attention in recent years. Palmprint recognition mainly focuses on image acquisition, preprocessing, feature selection and image matching. Feature extraction and matching are usually the most essential processes in palmprint recognition, and most of the research is based on feature selection and image matching, and many researchers use rich knowledge in machine learning and computer vision to solve these problems. In this paper, we propose a deep multi-view representation learning based multi-spectral palmprint fusion method, which uses deep neural networks to extract feature representation of multi-spectral palmprint images for palmprint classification. In this manner, the unique features of different spectral palmprint images can be used to learn a view-invariant representation of each palmprint. By using view-invariant representation, we can get better palmprint recognition performance than single modality. Experiments are performed on PolyU palmprint data set to validate the effectiveness of the proposed method.
Read moreModern Approaches to Software Vulnerability Detection: A Survey of Machine Learning, Deep Learning, and Large Language Models
Software vulnerabilities pose significant risks to the security and reliability of modern systems, making automated vulnerability detection an essential research area. Traditional static and rule-based approaches are limited in scalability and adaptability, motivating the adoption of data-driven methods. In this survey, we present a comprehensive review of Machine Learning (ML), Deep Learning (DL), and Large Language Models (LLMs) techniques for vulnerability detection. We analyze recent advances in feature representation, fine-tuning strategies, generative approaches, and prompt engineering, while highlighting their ability to capture both syntactic and semantic properties of source code. Furthermore, we examine commonly used evaluation metrics and provide a critical discussion of key challenges, including the lack of large-scale real-world datasets, limited vulnerability coverage, class imbalance, interpretability gaps, hallucination, and high computational costs. To address these issues, we outline promising future research directions, such as neuro-symbolic hybrid methods, parameter-efficient fine-tuning, continual learning, cross-language generalization, and explainable AI for vulnerability detection. Unlike previous studies, the present work explores learning paradigms from ML to LLMs using comprehensive evaluation criteria that highlight analytical capability, feature interpretability, and code-context comprehension. By combining these factors, our study addresses the methodological gap between classic feature-based approaches and current LLM-driven reasoning frameworks, providing beneficial insights to develop robust, scalable, and trustworthy software vulnerability detection systems.
Read moreDetecting Cyber Attacks in Smart Grids Using Semi-Supervised Anomaly Detection and Deep Representation Learning
Smart grids integrate advanced information and communication technologies (ICTs) into traditional power grids for more efficient and resilient power delivery and management, but also introduce new security vulnerabilities that can be exploited by adversaries to launch cyber attacks, causing severe consequences such as massive blackout and infrastructure damages. Existing machine learning-based methods for detecting cyber attacks in smart grids are mostly based on supervised learning, which need the instances of both normal and attack events for training. In addition, supervised learning requires that the training dataset includes representative instances of various types of attack events to train a good model, which is sometimes hard if not impossible. This paper presents a new method for detecting cyber attacks in smart grids using PMU data, which is based on semi-supervised anomaly detection and deep representation learning. Semi-supervised anomaly detection only employs the instances of normal events to train detection models, making it suitable for finding unknown attack events. A number of popular semi-supervised anomaly detection algorithms were investigated in our study using publicly available power system cyber attack datasets to identify the best-performing ones. The performance comparison with popular supervised algorithms demonstrates that semi-supervised algorithms are more capable of finding attack events than supervised algorithms. Our results also show that the performance of semi-supervised anomaly detection algorithms can be further improved by augmenting with deep representation learning.
Read moreCode Vulnerability Identification and Code Improvement using Advanced Machine Learning
Cyber-attacks are fairly mundane. The misconfigurations of the source code can result in security vulnerabilities that potentially encourage the attackers to exploit them and compromise the system. This paper aims to discover various mechanisms of automating the detection and correction of vulnerabilities in source code. Usage of static and dynamic analysis, various machine learning, deep learning, and neural network techniques will enhance the automation of detecting and correcting processes. This paper systematically presents the various methods and research efforts of detecting vulnerabilities in the source code, starting with what is a software vulnerability and what kind of exploitation, existing vulnerability detection methods, correction methods and efforts of best researches in the world relevant to the research area. A plugin will be developed which is capable of intelligently and efficiently detecting the vulnerable source code segment and correcting the source code accurately in the development stage.
Read moreVideo semantic segmentation using deep multi-view representation learning
In this paper, we propose a deep learning model based on deep multi-view representation learning, to address the video object segmentation task. The proposed model emphasizes the importance of the inherent correlation between video frames and incorporates a multi-view representation learning based on deep canonically correlated autoencoders. The multi-view representation learning in our model provides an efficient mechanism for capturing inherent correlations by jointly extracting useful features and learning better representation into a joint feature space, i.e., shared representation. To increase the training data and the learning capacity, we train the proposed model with pairs of video frames, i.e., Fa and Fb. During the segmentation phase, the deep canonically correlated autoencoders model encodes useful features by processing multiple reference frames together, which is used to detect the frequently reappearing. Our model enhances the state-of-the-art deep learning-based methods that mainly focus on learning discriminative foreground representations over appearance and motion. Experimental results over two large benchmarks demonstrate the ability of the proposed method to outperform competitive approaches and to reach good performances, in terms of semantic segmentation.
Read moreLDADEEP+: Latent aspect discovery with deep representations
Nowadays, with the success and fast growth of social media communities and mobile devices, people are encouraged to share their multimedia data online. Analyzing and summarizing data into useful information thus becomes increasingly important. For online photo sharing services like Flickr, when users are uploading a batch of daily photos at a time, the tags users provided tend to be rather vague, containing only a small amount of information. For better photo application and understanding, we attempt to automatically discover semantic-rich (hidden) aspects of photos merely by looking at image contents. In this paper, we propose an effective model, which is a combination of LDA model and deep learning representations, to realize the idea of automatic aspect discovery. We then discuss the properties of this aspect discovery model through experiments on event summarization task. In those experiments, we show the high diversity and high quality of aspects discovered by our proposed method. Meanwhile, we conduct an user study to evaluate the quality of the summarized results. Moreover, the proposed method can be further extended to human attribute discovery for a given event. We automatically discover different aspects on our Olympic Games data (e.g. football, ice skating).
Read moreA transformer-based framework for software vulnerability detection using attention-driven convolutional neural networks
A transformer-based framework for software vulnerability detection using attention-driven convolutional neural networks
AIMAFE: Autism spectrum disorder identification with multi-atlas deep feature representation and ensemble learning
AIMAFE: Autism spectrum disorder identification with multi-atlas deep feature representation and ensemble learning