- Research Article
3
- 10.1016/j.patcog.2023.109976
Learning node representations against perturbations
- Sep 18, 2023
- Pattern Recognition
- Xu Chen + 3 more +3
Learning node representations against perturbations
Graph neural networks (GNNs) have been widely used to learn node representations from graph data in an unsupervised way for downstream tasks. However, when applied to detect anomalies (e.g., outliers, unexpected density), they deliver unsatisfactory performance as existing loss functions fail. For example, any loss based on random walk (RW) algorithms would no longer work because the assumption that anomalous nodes were close with each other could not hold. Moreover, the nature of class imbalance in anomaly detection tasks brings great challenges to reduce the prediction error. In this work, we propose a novel loss function to train GNNs for anomaly-detectable node representations. It evaluates node similarity using global grouping patterns discovered from graph mining algorithms. It can automatically adjust margins for minority classes based on data distribution. Theoretically, we prove that the prediction error is bounded given the proposed loss function. We empirically investigate the GNN effectiveness of different loss variants based on different algorithms. Experiments on two real-world datasets show that they perform significantly better than RW-based loss for graph anomaly detection.
Learning node representations against perturbations
Learning node representations against perturbations
Learning and Reasoning with Graph Data: Neural and Statistical-Relational Approaches (Invited Paper)
Graph neural networks (GNNs) have emerged in recent years as a very powerful and popular modeling tool for graph and network data. Though much of the work on GNNs has focused on graphs with a single edge relation, they have also been adapted to multi-relational graphs, including knowledge graphs. In such multi-relational domains, the objectives and possible applications of GNNs become quite similar to what for many years has been investigated and developed in the field of statistical relational learning (SRL). This article first gives a brief overview of the main features of GNN and SRL approaches to learning and reasoning with graph data. It analyzes then in more detail their commonalities and differences with respect to semantics, representation, parameterization, interpretability, and flexibility. A particular focus will be on relational Bayesian networks (RBNs) as the SRL framework that is most closely related to GNNs. We show how common GNN architectures can be directly encoded as RBNs, thus enabling the direct integration of "low level" neural model components with the "high level" symbolic representation and flexible inference capabilities of SRL.
Read moreIDEA: A Flexible Framework of Certified Unlearning for Graph Neural Networks
Graph Neural Networks (GNNs) have been increasingly deployed in a plethora of applications. However, the graph data used for training may contain sensitive personal information of the involved individuals. Once trained, GNNs typically encode such information in their learnable parameters. As a consequence, privacy leakage may happen when the trained GNNs are deployed and exposed to potential attackers. Facing such a threat, machine unlearning for GNNs has become an emerging technique that aims to remove certain personal information from a trained GNN. Among these techniques, certified unlearning stands out, as it provides a solid theoretical guarantee of the information removal effectiveness. Nevertheless, most of the existing certified unlearning methods for GNNs are only designed to handle node and edge unlearning requests. Meanwhile, these approaches are usually tailored for either a specific design of GNN or a specially designed training objective. These disadvantages significantly jeopardize their flexibility. In this paper, we propose a principled framework named IDEA to achieve flexible and certified unlearning for GNNs. Specifically, we first instantiate four types of unlearning requests on graphs, and then we propose an approximation approach to flexibly handle these unlearning requests over diverse GNNs. We further provide theoretical guarantee of the effectiveness for the proposed approach as a certification. Different from existing alternatives, IDEA is not designed for any specific GNNs or optimization objectives to perform certified unlearning, and thus can be easily generalized. Extensive experiments on real-world datasets demonstrate the superiority of IDEA in multiple key perspectives.
Read moreLGNN: a novel linear graph neural network algorithm.
The emergence of deep learning has not only brought great changes in the field of image recognition, but also achieved excellent node classification performance in graph neural networks. However, the existing graph neural network framework often uses methods based on spatial domain or spectral domain to capture network structure features. This process captures the local structural characteristics of graph data, and the convolution process has a large amount of calculation. It is necessary to use multi-channel or deep neural network structure to achieve the goal of modeling the high-order structural characteristics of the network. Therefore, this paper proposes a linear graph neural network framework [Linear Graph Neural Network (LGNN)] with superior performance. The model first preprocesses the input graph, and uses symmetric normalization and feature normalization to remove deviations in the structure and features. Then, by designing a high-order adjacency matrix propagation mechanism, LGNN enables nodes to iteratively aggregate and learn the feature information of high-order neighbors. After obtaining the node representation of the network structure, LGNN uses a simple linear mapping to maintain computational efficiency and obtain the final node representation. The experimental results show that the performance of the LGNN algorithm in some tasks is slightly worse than that of the existing mainstream graph neural network algorithms, but it shows or exceeds the machine learning performance of the existing algorithms in most graph neural network performance evaluation tasks, especially on sparse networks.
Read morePower System Network Topology Identification Based on Knowledge Graph and Graph Neural Network
The automatic identification of the topology of power networks is important for the data-driven and situation-aware operation of power grids. Traditional methods of topology identification lack a data-tolerant mechanism, and the accuracy of their performance in terms of identification is thus affected by the quality of data. Topology identification is related to the link prediction problem. The graph neural network can be used to predict the state of unlabeled nodes (lines) through training on features of labeled nodes (lines) with fault tolerance. Inspired by the characteristics of the graph neural network, we applied it to topology identification in this study. We propose a method to identify the topology of a power network based on a knowledge graph and the graph neural network. Traditional knowledge graphs can quickly mine possible connections between entities and generate graph structure data, but in the case of errors or informational conflicts in the data, they cannot accurately determine whether the relationships between the entities really exist. The graph neural network can use data mining to determine whether a connection obtained between entities is based on their eigenvalues, and has a fault tolerance mechanism to adapt to errors and informational conflicts in the graph data, but needs the graph data as database. The combination of the knowledge graph and the graph neural network can compensate for the deficiency of the single knowledge graph method. We tested the proposed method by using the IEEE 118-bus system and a provincial network system. The results showed that our approach is feasible and highly fault tolerant. It can accurately identify network topology even in the presence of conflicting and missing measurement-related information.
Read moreGraph vector function architecture.
Graph Neural Networks (GNNs) are the most common approach for learning complex relational data represented using graph data structures. Although GNNs are effective at learning representations of both nodes and graphs for a given task, the learning process is computationally expensive and as such, time and energy-inefficient. This paper investigates this challenge within the context of recent work on untrained graph representations that only train the solver model. We present Graph Vector Function Architecture (GVFA), a novel alternative to learning graph representations in GNNs that is based on hyperdimensional computing (HDC) principles. GVFA is a general zero-shot approach for graph and node representations without learning. As such, our representations are not task-specific and the computational costs of constructing them is substantially lower compared to learning-based GNN. Empirically, we demonstrate the expressiveness and generalization properties of different GVFA configurations. Our experimental results demonstrate that GVFA outperforms several classic GNNs on their benchmark datasets in terms of classification accuracy for both graph and node classification tasks, while also yielding a substantial reduction in training time.
Read moreNode Classification in GNNs: Impact of Neighborhood Label Distribution on Homophily and Heterophily.
In node classification, traditional graph neural networks (GNNs) typically assume implicit homophily, indicating that intraclass nodes are likely connected. However, real-world graphs frequently exhibit heterophily, in which interclass nodes are also commonly connected. To address this challenge, recent methods have adopted approaches such as expanding local neighborhoods and employing adaptive message aggregation to enhance the GNN performance on heterophily graphs. Nevertheless, these methods are restricted by the homophily assumption and fail to effectively capture long-range dependencies (e.g., widely separated intraclass nodes) and insufficiently leverage the graph topology. This study investigates the performance differences of GNN when it is applied to both homophily and heterophily graphs and finds that the distinguishability of neighborhood label distributions (NLDs) exhibits a significant correlation with the accuracy of node classification. To assess the impact of NLD on node classification, this study proposes a novel homophily metric based on node distinguishability. Subsequently, this study introduces a new GNN model named NLD-based GNN (NLDGNN) for node classification. First, NLDGNN initializes node representations by integrating node features with node NLDs. To address long-range dependencies in heterophily graphs, NLDGNN utilizes the global label relationship matrix with low-rank characteristics for global message passing. By combining the attention scores derived from the initial node representations, NLDGNN constructs the global label relationship matrix for enhanced message passing, thereby improving the expressiveness of node representations. Experimental results indicate that NLDGNN outperforms existing GNN models on both real-world homophily and heterophily graphs. The code of this study is available at https://github.com/wanli6/NLDGNN.
Read moreStructack: Structure-based Adversarial Attacks on Graph Neural Networks
Recent work has shown that graph neural networks (GNNs) are vulnerable to adversarial attacks on graph data. Common attack approaches are typically informed, i.e. they have access to information about node attributes such as labels and feature vectors. In this work, we study adversarial attacks that are uninformed, where an attacker only has access to the graph structure, but no information about node attributes. Here the attacker aims to exploit structural knowledge and assumptions, which GNN models make about graph data. In particular, literature has shown that structural node centrality and similarity have a strong influence on learning with GNNs. Therefore, we study the impact of centrality and similarity on adversarial attacks on GNNs. We demonstrate that attackers can exploit this information to decrease the performance of GNNs by focusing on injecting links between nodes of low similarity and, surprisingly, low centrality. We show that structure-based uninformed attacks can approach the performance of informed attacks, while being computationally more efficient. With our paper, we present a new attack strategy on GNNs that we refer to as Structack. Structack can successfully manipulate the performance of GNNs with very limited information while operating under tight computational constraints. Our work contributes towards building more robust machine learning approaches on graphs.
Read moreFair Graph Representation Learning via Diverse Mixture-of-Experts
Graph Neural Networks (GNNs) have demonstrated a great representation learning capability on graph data and have been utilized in various downstream applications. However, real-world data in web-based applications (e.g., recommendation and advertising) always contains bias, preventing GNNs from learning fair representations. Although many works were proposed to address the fairness issue, they suffer from the significant problem of insufficient learnable knowledge with limited attributes after debiasing. To address this problem, we develop Graph-Fairness Mixture of Experts (G-Fame), a novel plug-and-play method to assist any GNNs to learn distinguishable representations with unbiased attributes. Furthermore, based on G-Fame, we propose G-Fame++, which introduces three novel strategies to improve the representation fairness from node representations, model layer, and parameter redundancy perspectives. In particular, we first present the embedding diversified method to learn distinguishable node representations. Second, we design the layer diversified strategy to maximize the output difference of distinct model layers. Third, we introduce the expert diversified method to minimize expert parameter similarities to learn diverse and complementary representations. Extensive experiments demonstrate the superiority of G-Fame and G-Fame++ in both accuracy and fairness, compared to state-of-the-art methods across multiple graph datasets.
Read moreGraph Neural Network Learning on the Pediatric Structural Connectome.
Sex classification is a major benchmark of previous work in learning on the structural connectome, a naturally occurring brain graph that has proven useful for studying cognitive function and impairment. While graph neural networks (GNNs), specifically graph convolutional networks (GCNs), have gained popularity lately for their effectiveness in learning on graph data, achieving strong performance in adult sex classification tasks, their application to pediatric populations remains unexplored. We seek to characterize the capacity for GNN models to learn connectomic patterns on pediatric data through an exploration of training techniques and architectural design choices. Two datasets comprising an adult BRIGHT dataset (N = 147 Hodgkin's lymphoma survivors and N = 162 age similar controls) and a pediatric Human Connectome Project in Development (HCP-D) dataset (N = 135 healthy subjects) were utilized. Two GNN models (GCN simple and GCN residual), a deep neural network (multi-layer perceptron), and two standard machine learning models (random forest and support vector machine) were trained. Architecture exploration experiments were conducted to evaluate the impact of network depth, pooling techniques, and skip connections on the ability of GNN models to capture connectomic patterns. Models were assessed across a range of metrics including accuracy, AUC score, and adversarial robustness. GNNs outperformed other models across both populations. Notably, adult GNN models achieved 85.1% accuracy in sex classification on unseen adult participants, consistent with prior studies. The extension of the adult models to the pediatric dataset and training on the smaller pediatric dataset were sub-optimal in their performance. Using adult data to augment pediatric models, the best GNN achieved comparable accuracy across unseen pediatric (83.0%) and adult (81.3%) participants. Adversarial sensitivity experiments showed that the simple GCN remained the most robust to perturbations, followed by the multi-layer perceptron and the residual GCN. These findings underscore the potential of GNNs in advancing our understanding of sex-specific neurological development and disorders and highlight the importance of data augmentation in overcoming challenges associated with small pediatric datasets. Further, they highlight relevant tradeoffs in the design landscape of connectomic GNNs. For example, while the simpler GNN model tested exhibits marginally worse accuracy and AUC scores in comparison to the more complex residual GNN, it demonstrates a higher degree of adversarial robustness.
Read moreMIGP: Metapath Integrated Graph Prompt Neural Network
MIGP: Metapath Integrated Graph Prompt Neural Network
KNN-GNN: A powerful graph neural network enhanced by aggregating K-nearest neighbors in common subspace
KNN-GNN: A powerful graph neural network enhanced by aggregating K-nearest neighbors in common subspace
Lotan: Bridging the Gap between GNNs and Scalable Graph Analytics Engines
Recent advances in Graph Neural Networks (GNNs) have changed the landscape of modern graph analytics. The complexity of GNN training and the scalability challenges have also sparked interest from the systems community, with efforts to build systems that provide higher efficiency and schemes to reduce costs. However, we observe that many such systems basically "reinvent the wheel" of much work done in the database world on scalable graph analytics engines. Further, they often tightly couple the scalability treatments of graph data processing with that of GNN training, resulting in entangled complex problems and systems that often do not scale well on one of those axes. In this paper, we ask a fundamental question: How far can we push existing systems for scalable graph analytics and deep learning (DL) instead of building custom GNN systems? Are compromises inevitable on scalability and/or runtimes? We propose Lotan, the first scalable and optimized data system for full-batch GNN training with decoupled scaling that bridges the hitherto siloed worlds of graph analytics systems and DL systems. Lotan offers a series of technical innovations, including re-imagining GNN training as query plan-like dataflows, execution plan rewriting, optimized data movement between systems, a GNN-centric graph partitioning scheme, and the first known GNN model batching scheme. We prototyped Lotan on top of GraphX and PyTorch. An empirical evaluation using several real-world benchmark GNN workloads reveals a promising nuanced picture: Lotan significantly surpasses the scalability of state-of-the-art custom GNN systems, while often matching or being only slightly behind on time-to-accuracy metrics in some cases. We also show the impact of our system optimizations. Overall, our work shows that the GNN world can indeed benefit from building on top of scalable graph analytics engines. Lotan's new level of scalability can also empower new ML-oriented research on ever-larger graphs and GNNs.
Read moreKnowledge Enhanced Graph Neural Networks
Graph data is omnipresent and has a wide variety of applications, such as in natural science, social networks, or the semantic web. However, while being rich in information, graphs are often noisy and incomplete. As a result, graph completion tasks, such as node classification or link prediction, have gained attention. On one hand, neural methods, such as graph neural networks, have proven to be robust tools for learning rich representations of noisy graphs. On the other hand, symbolic methods enable exact reasoning on graphs. We propose Knowledge Enhanced Graph Neural Networks (KeGNN), a neuro-symbolic framework for graph completion that combines both paradigms as it allows for the integration of prior knowledge into a graph neural network model. Essentially, KeGNN consists of a graph neural network as a base upon which knowledge enhancement layers are stacked with the goal of refining predictions with respect to prior knowledge. We instantiate KeGNN in conjunction with two well-known graph neural networks, Graph Convolutional Networks and Graph Attention Networks, and evaluate KeGNN on multiple benchmark datasets for node classification.
Read moreHypergraph Link Prediction: Learning Drug Interaction Networks Embeddings
Graph neural networks (GNNs) have revolutionized deep learning on non-Euclidean data domains, and are extensively used in fields such as social media and recommendation systems. However, complex relational data structures such as hypergraphs, pose challenges for GNNs in terms of their ability to model, embed, and learn relational complexities of multigraphs. Most GNNs focus on capturing flat local neighborhoods of a node thus failing to account for structural properties of multi-relational graphs. This paper introduces Hypergraph Link Prediction (HLP), a novel approach of encoding the multilink structure of graphs. HLP allows pooling operations to incorporate a 360 degrees overview of a node interaction profile, by learning local neighborhood and global hypergraph structure simultaneously. Global graph information is injected into node representations, such that unique global structural patterns of every node are encoded at the node level. HLP leverages the augmented hypergraph adjacency matrix to incorporate the depth of the hypergraph in the convolutional layers. The model is applied to the task of predicting multi-drug interactions, by modeling relations between pairs of drugs as a hypergraph. The existence and the type of drug interactions between the same pair of drugs are mapped as multiple edges, and can be inferred by learning the multigraph local and global structure concurrently. To account for molecular graph properties of a drug, additional drug chemical graph structural fingerprints are included as node attributes.
Read more