- Research Article
3
- 10.1016/j.patcog.2023.109976
Learning node representations against perturbations
- Sep 18, 2023
- Pattern Recognition
- Xu Chen + 3 more +3
Learning node representations against perturbations
Node representation learning on attributed graphs-whose nodes are associated with rich attributes (e.g., texts and protein sequences)-plays a crucial role in many important downstream tasks. To encode the attributes and graph structures simultaneously, recent studies integrate pre-trained models with graph neural networks (GNNs), where pre-trained models serve as node encoders (NEs) to encode the attributes. As jointly training large NEs and GNNs on large-scale graphs suffers from severe scalability issues, many methods propose to train NEs and GNNs separately. Consequently, they do not take feature convolutions in GNNs into consideration in the training phase of NEs, leading to a significant learning bias relative to the joint training. To address this challenge, we propose an efficient label regularization technique, namely Label Deconvolution (LD), to alleviate the learning bias by a novel and highly scalable approximation to the inverse mapping of GNNs. The inverse mapping leads to an objective function that is equivalent to that by the joint training, while it can effectively incorporate GNNs in the training phase of NEs against the learning bias. More importantly, we show that LD converges to the optimal objective function values by the joint training under mild assumptions. Experiments demonstrate LD significantly outperforms state-of-the-art methods on Open Graph Benchmark datasets.
Learning node representations against perturbations
Learning node representations against perturbations
Why are Graph Neural Networks Effective for EDA Problems?
In this paper, we discuss the source of effectiveness of Graph Neural Networks (GNNs) in EDA, particularly in the VLSI design automation domain. We argue that the effectiveness comes from the fact that GNNs implicitly embed the prior knowledge and inductive biases associated with given VLSI tasks, which is one of the three approaches to make a learning algorithm physics-informed. These inductive biases are different to those common used in GNNs designed for other structured data, such as social networks and citation networks. We will illustrate this principle with several recent GNN examples in the VLSI domain, including predictive tasks such as switching activity prediction, timing prediction, parasitics prediction, layout symmetry prediction, as well as optimization tasks such as gate sizing and macro and cell transistor placement. We will also discuss the challenges of applications of GNN and the opportunity of applying self-supervised learning techniques with GNN for VLSI optimization.
Read moreGraph Representation Learning with Adaptive Mixtures
Graph Neural Networks (GNNs) are the current state-of-the-art models in learning node representations for many predictive tasks on graphs. Typically, GNNs reuses the same set of model parameters across all nodes in the graph to improve the training efficiency and exploit the translationally-invariant properties in many datasets. However, the parameter sharing scheme prevents GNNs from distinguishing two nodes that are isomorphic and that the translation invariance property may not exhibit in real-world graphs. In this paper, we present Graph Representation Learning with Adaptive Mixtures (GRAM), a novel approach for learning node representations in a graph by introducing multiple independent GNN models and a trainable mixture distribution for each node. GRAM is a general framework that can be readily embedded into existing models for enhancing representation learning. It allows for adaptive node clustering with mixtures and joint optimization of all GNN models simultaneously. To achieve this, we develop an Expectation-Maximization algorithm with stochastic gradient descent in a scalable manner. Specifically, in the E-Step, we approximate the posterior probability distribution of the latent cluster membership based on the k-th independent GNN models in the last iteration. In the M-step, we update all GNN models and the mixture parameters based on the refined latent cluster membership for downstream tasks, e.g., node classification. To fully exploit the graph structural information, we further implement a regularization term on unlabeled nodes, which encourages adjacent nodes to have the same mixture distribution. By jointly optimizing the model in an end-to-end manner, GRAM learns to capture different abstract patterns on graphs for different node clusters, which yields discriminative representation for each node. We evaluate GRAM on five benchmark datasets with extensive experiments. GRAM is demonstrated to consistently boost state-of-the-art GNN variants in node classification tasks.
Read moreKNN-GNN: A powerful graph neural network enhanced by aggregating K-nearest neighbors in common subspace
KNN-GNN: A powerful graph neural network enhanced by aggregating K-nearest neighbors in common subspace
AI-Driven molecule generation and bioactivity prediction: A multi-model approach combining VAE, graph and language-based neural networks.
AI-Driven molecule generation and bioactivity prediction: A multi-model approach combining VAE, graph and language-based neural networks.
Read moreGeometric deep learning for online prediction of cascading failures in power grids
Past events have revealed that widespread blackouts are mostly a result of cascading failures in the power grid. Understanding the underlining mechanisms of cascading failures can help in developing strategies to minimize the risk of such events. Moreover, a real-time detection of precursors to cascading failures will help operators take measures to prevent their propagation. Currently, the well-established probabilistic and physics-based models of cascading failures offer low computational efficiency, hindering them to be used only as offline tools. In this work, we develop a data-driven methodology for online estimation of the risk of cascading failures. We utilize a physics-based cascading failure model to generate a cascading failure dataset considering different operating conditions and failure scenarios, thus obtaining a sample space covering a large set of power grid states that are labeled as safe or unsafe. We use the synthetic data to train deep learning architectures, namely Feed-forward Neural Networks (FNN) and Graph Neural Networks (GNN). With the development of GNNs, improved performance is achieved with graph-structured data, and GNNs can generalize to graphs of diverse sizes. A comparison between FNN and GNN is made and the GNNs inductive capability is demonstrated via test grids. Furthermore, we apply transfer learning to improve the performance of a pre-trained GNN model on power grids not seen in the training process. The GNN model shows accuracy and balanced accuracy above 96% on selected test datasets not used in the training. Conversely, the FNN shows accuracy above 85% and balanced accuracy above 81% on test datasets unseen during training. Overall, the GNN model is successful in determining, if one or several simultaneous outages result in a critical grid state, under specific grid operating conditions.
Read moreHierarchical message-passing graph neural networks
Graph Neural Networks (GNNs) have become a prominent approach to machine learning with graphs and have been increasingly applied in a multitude of domains. Nevertheless, since most existing GNN models are based on flat message-passing mechanisms, two limitations need to be tackled: (i) they are costly in encoding long-range information spanning the graph structure; (ii) they are failing to encode features in the high-order neighbourhood in the graphs as they only perform information aggregation across the observed edges in the original graph. To deal with these two issues, we propose a novel Hierarchical Message-passing Graph Neural Networks framework. The key idea is generating a hierarchical structure that re-organises all nodes in a flat graph into multi-level super graphs, along with innovative intra- and inter-level propagation manners. The derived hierarchy creates shortcuts connecting far-away nodes so that informative long-range interactions can be efficiently accessed via message passing and incorporates meso- and macro-level semantics into the learned node representations. We present the first model to implement this framework, termed Hierarchical Community-aware Graph Neural Network (HC-GNN), with the assistance of a hierarchical community detection algorithm. The theoretical analysis illustrates HC-GNN’s remarkable capacity in capturing long-range information without introducing heavy additional computation complexity. Empirical experiments conducted on 9 datasets under transductive, inductive, and few-shot settings exhibit that HC-GNN can outperform state-of-the-art GNN models in network analysis tasks, including node classification, link prediction, and community detection. Moreover, the model analysis further demonstrates HC-GNN’s robustness facing graph sparsity and the flexibility in incorporating different GNN encoders.
Read moreAn integrated graph data privacy attack framework based on graph neural networks in IoT
SummaryKnowledge graphs contain a large amount of entity and relational data, and graph neural networks, as a class of efficient graph representation techniques based on deep learning, excel in knowledge graph modeling. However, previous neural network architectures for the most part only learn node representations and do not fully consider the heterogeneity of data. In this article, we innovatively propose a privacy attack framework based on IoT, PAFI, which is able to classify entities and relations, learn embedding representations in multi‐relational graphs, and can be applied to some existing neural network algorithms. Based on this, a fine‐grained privacy attack model, FPM, is proposed, which can perform attack operations on multiple targets, achieve selectivity of target tasks, and greatly improve the generalization ability of the attack model. In this article, the effectiveness of PAFI and FPM is demonstrated by real network datasets, and compared with previous attack methods, both of which achieve good results.
Read moreSALE-MLP: Structure Aware Latent Embeddings for GNN to Graph-free MLP Distillation
Graph Neural Networks (GNNs), with their ability to effectively handle non-Euclidean data structures, have demonstrated state-of-the-art performance in learning node and graph-level representations. However, GNNs face significant computational overhead due to their message-passing mechanisms, making them impractical for real-time large-scale applications. Recently, Graph-to-MLP (G2M) knowledge distillation has emerged as a promising solution, utilizing MLPs to reduce inference latency. However, existing methods often lack structural awareness (SA), limiting their ability to capture essential graph-specific information. Moreover, some methods require access to large-scale graphs, undermining their scalability. To address these issues, we propose SALE-MLP (Structure-Aware Latent Embeddings for GNN-to-Graph-Free MLP Distillation), a novel graph-free and structure-aware approach that leverages unsupervised structural losses to align the MLP feature space with the underlying graph structure. SALE-MLP does not rely on precomputed GNN embeddings nor require graph during inference, making it efficient for real-world applications. Extensive experiments demonstrate that SALE-MLP outperforms existing G2M methods across tasks and datasets, achieving 3–4% improvement in node classification for inductive settings while maintaining strong transductive performance.
Read moreLong Short-Term Graph Memory Against Class-imbalanced Over-smoothing
Most Graph Neural Networks (GNNs) follow the message-passing scheme. Residual connection is an effective strategy to tackle GNNs' over-smoothing issue and performance reduction issue on non-homophilic networks. Unfortunately, the coarse-grained residual connection still suffers from class-imbalanced over-smoothing issue, due to the fixed and linear combination of topology and attribute in node representation learning. To make the combination flexible to capture complicated relationship, this paper reveals that the residual connection needs to be node-dependent, layer-dependent, and related to both topology and attribute. To alleviate the difficulty in specifying complicated relationship, this paper presents a novel perspective on GNNs, i.e., the representations of one node in different layers can be seen as a sequence of states. From this perspective, existing residual connections are not flexible enough for sequence modeling. Therefore, a novel node-dependent residual connection, i.e., Long Short-Term Graph Memory Network (LSTGM) is proposed to employ Long Short-Term Memory (LSTM), to model the sequence of node representation. To make the graph topology fully employed, LSTGM innovatively enhances the updated memory and three gates with graph topology. A speedup version is also proposed for effective training. Experimental evaluations on real-world datasets demonstrate their effectiveness in preventing over-smoothing issue and handling networks with heterophily.
Read moreInteraction-Based Inductive Bias in Graph Neural Networks: Enhancing Protein-Ligand Binding Affinity Predictions From 3D Structures.
Inductive bias in machine learning (ML) is the set of assumptions describing how a model makes predictions. Different ML-based methods for protein-ligand binding affinity (PLA) prediction have different inductive biases, leading to different levels of generalization capability and interpretability. Intuitively, the inductive bias of an ML-based model for PLA prediction should fit in with biological mechanisms relevant for binding to achieve good predictions with meaningful reasons. To this end, we propose an interaction-based inductive bias to restrict neural networks to functions relevant for binding with two assumptions: 1) A protein-ligand complex can be naturally expressed as a heterogeneous graph with covalent and non-covalent interactions; 2) The predicted PLA is the sum of pairwise atom-atom affinities determined by non-covalent interactions. The interaction-based inductive bias is embodied by an explainable heterogeneous interaction graph neural network (EHIGN) for explicitly modeling pairwise atom-atom interactions to predict PLA from 3D structures. Extensive experiments demonstrate that EHIGN achieves better generalization capability than other state-of-the-art ML-based baselines in PLA prediction and structure-based virtual screening. More importantly, comprehensive analyses of distance-affinity, pose-affinity, and substructure-affinity relations suggest that the interaction-based inductive bias can guide the model to learn atomic interactions that are consistent with physical reality. As a case study to demonstrate practical usefulness, our method is tested for predicting the efficacy of Nirmatrelvir against SARS-CoV-2 variants. EHIGN successfully recognizes the changes in the efficacy of Nirmatrelvir for different SARS-CoV-2 variants with meaningful reasons.
Read moreAS-GCN: Adaptive Semantic Architecture of Graph Convolutional Networks for Text-Rich Networks
Graph Neural Networks (GNNs) have demonstrated great power in many network analytical tasks. However, graphs (i.e., networks) in the real world are usually text-rich, implying that valuable semantic information needs to be carefully considered. Existing GNNs for text-rich networks typically treat text as attribute words alone, which inevitably leads to the loss of important semantic structures, limiting the representation capability of GNNs. In this paper, we propose an end-to-end adaptive semantic architecture of graph convolutional networks, namely AS-GCN, which unifies neural topic model and graph convolutional networks, for text-rich network representation. Specifically, we utilize a neural topic model to extract the global topic semantics, and accordingly augment the original text-rich network into a tri-typed heterogeneous network, capturing both the local word-sequence semantic structure and the global topic semantic structure from text. We then design an effective semantic-aware propagation of information by introducing a discriminative convolution mechanism. We further propose two strategies, that is, distribution sharing and joint training, to adaptively generate a proper network structure based on the learning objective to improve network representation. Extensive experiments on text-rich networks illustrate that our new architecture outperforms the state-of-the-art methods by a significant improvement. Meanwhile, this architecture can also be applied to e-commerce search scenes, and experiments on a real e-commerce problem from JD further demonstrate the superiority of the proposed architecture over the baselines.
Read moreTransferable surrogate models based on inductive biases of graph neural networks for water distribution systems
<p>Water utilities tackle various problems in planning and operating their systems with complex and computationally expensive hydraulic models, i.e., maximizing system resilience, fault isolation, risk assessment, optimal pump scheduling, or water loss reduction via pressure management. To meet limited computational budgets, engineers employ less resource-intensive surrogate models. Current surrogate models based on artificial neural networks deliver similar accuracies as hydraulic models with lower computational costs. However, they require retraining when applied to an unknown water distribution system, which increases their computational load and limits their general applicability. Recent advancements in graph-based machine learning address these limitations. Graph neural networks (GNNs) naturally connect with the network elements (e.g., pipes and valves with edges, junctions, and tanks with vertices)  of water distribution systems, proving themselves to be a promising candidate for surrogate modeling. Once trained on a specific network to be a surrogate model, GNNs possess inductive biases that allow transferability to an unseen topology. In this work, we adopted a demand-driven simulation of a water distribution system in a graph machine learning setting. We built a synthetic dataset of demand-driven simulation with EPANET, founded on the example of real-world water distribution systems, and trained an attention-based GNN to emulate the hydraulic simulator. The accuracy was evaluated inductively on an unseen larger-sized water distribution network. We observed that the model showed promising transferability results to a larger network without the need for additional re-training on the unseen topology.</p>
Read morePre-Training Graph Neural Networks for Cold-Start Users and Items Representation
Cold-start problem is a fundamental challenge for recommendation tasks. Despite the recent advances on Graph Neural Networks (GNNs) incorporate the high-order collaborative signal to alleviate the problem, the embeddings of the cold-start users and items aren't explicitly optimized, and the cold-start neighbors are not dealt with during the graph convolution in GNNs. This paper proposes to pre-train a GNN model before applying it for recommendation. Unlike the goal of recommendation, the pre-training GNN simulates the cold-start scenarios from the users/items with sufficient interactions and takes the embedding reconstruction as the pretext task, such that it can directly improve the embedding quality and can be easily adapted to the new cold-start users/items. To further reduce the impact from the cold-start neighbors, we incorporate a self-attention-based meta aggregator to enhance the aggregation ability of each graph convolution step, and an adaptive neighbor sampler to select the effective neighbors according to the feedbacks from the pre-training GNN model. Experiments on three public recommendation datasets show the superiority of our pre-training GNN model against the original GNN models on user/item embedding inference and the recommendation task.
Read morePre-training Language Model for Friend Recommendation: A Case Study of Large Social Graph
Despite the demonstrated success of Pre-trained Language Models (PLM) in enhancing various recommendation tasks, their performance in friend recommendation systems, which are heavily influenced by complex social network dynamics, remains unproven. On the other hand, for social networks like Facebook, graph-based machine learning approaches in friend suggestion, like Graph Neural Networks (GNNs), struggle with the computational demands posed by social graphs. In is paper we investigate the application of PLM and GNN in the context of PYMK features within Facebook's social graph. We propose a representation learning scheme that captures both the local structure and semantic content within a second-degree connection space. Our approach leverages pre-trained transformer models with a feature aggregation scheme that enables efficient node representation learning in large social networks. We detail our model's architecture and discuss our training methods, accompanied by detailed experiments. Our model has improved the amount of friending on Facebook, facilitating new connections for users each day in online experiment.
Read more