- Research Article
63
- 10.1016/j.cosrev.2020.100296
Network embedding: Taxonomies, frameworks and applications
- Aug 26, 2020
- Computer Science Review
- Mingliang Hou + 5 more +5
Network embedding: Taxonomies, frameworks and applications
Network Embedding, which represents nodes in networks with efficient low-dimensional vectors, has been proved useful in a variety of applications. However, most existing approaches study single-view networks but not the multi-view networks with multiple types of relationships between nodes. Meanwhile, they ignore the rich features associated with the nodes, which is common in real world. In this paper, we propose a novel network embedding method, Intra-view and Inter-view attention for Multi-view Network Embedding (I2MNE), which leverages both the multi-view network structure and the node features to efficiently generate node representations. Specially, we introduce the intra-view attention when aggregating node features from neighbors for each single view and the inter-view attention when integrating representations across different views. Experiments on two real-world networks show that our approach outperforms other counterpart network embedding methods.
Network embedding: Taxonomies, frameworks and applications
Network embedding: Taxonomies, frameworks and applications
Incorporating Network Embedding into Markov Random Field for Better Community Detection
Recent research on community detection focuses on learning representations of nodes using different network embedding methods, and then feeding them as normal features to clustering algorithms. However, we find that though one may have good results by direct clustering based on such network embedding features, there is ample room for improvement. More seriously, in many real networks, some statisticallysignificant nodes which play pivotal roles are often divided into incorrect communities using network embedding methods. This is because while some distance measures are used to capture the spatial relationship between nodes by embedding, the nodes after mapping to feature vectors are essentially not coupled any more, losing important structural information. To address this problem, we propose a general Markov Random Field (MRF) framework to incorporate coupling in network embedding which allows better detecting network communities. By smartly utilizing properties of MRF, the new framework not only preserves the advantages of network embedding (e.g. low complexity, high parallelizability and applicability for traditional machine learning), but also alleviates its core drawback of inadequate representations of dependencies via making up the missing coupling relationships. Experiments on real networks show that our new approach improves the accuracy of existing embedding methods (e.g. Node2Vec, DeepWalk and MNMF), and corrects most wrongly-divided statistically-significant nodes, which makes network embedding essentially suitable for real community detection applications. The new approach also outperforms other state-of-the-art conventional community detection methods.
Read moreA Survey on Network Embedding
Network embedding assigns nodes in a network to low-dimensional representations and effectively preserves the network structure. Recently, a significant amount of progresses have been made toward this emerging network analysis paradigm. In this survey, we focus on categorizing and then reviewing the current development on network embedding methods, and point out its future research directions. We first summarize the motivation of network embedding. We discuss the classical graph embedding algorithms and their relationship with network embedding. Afterwards and primarily, we provide a comprehensive overview of a large number of network embedding methods in a systematic manner, covering the structure- and property-preserving network embedding methods, the network embedding methods with side information and the advanced information preserving network embedding methods. Moreover, several evaluation approaches for network embedding and some useful online resources, including the network data sets and softwares, are reviewed, too. Finally, we discuss the framework of exploiting these network embedding methods to build an effective system and point out some potential future directions.
Read moreAdversarial network embedding using structural similarity
Network embedding which aims to embed a given network into a low-dimensional vector space has been proved effective in various network analysis and mining tasks such as node classification, link prediction and network visualization. The emerging network embedding methods have shifted of emphasis in utilizing mature deep learning models. The neural-network based network embedding has become a mainstream solution because of its high efficiency and capability of preserving the nonlinear characteristics of the network. In this paper, we propose Adversarial Network Embedding using Structural Similarity (ANESS), a novel, versatile, low-complexity GAN-based network embedding model which utilizes the inherent vertex-to-vertex structural similarity attribute of the network. ANESS learns robustness and effective vertex embeddings via a adversarial training procedure. Specifically, our method aims to exploit the strengths of generative adversarial networks in generating high-quality samples and utilize the structural similarity identity of vertexes to learn the latent representations of a network. Meanwhile, ANESS can dynamically update the strategy of generating samples during each training iteration. The extensive experiments have been conducted on the several benchmark network datasets, and empirical results demonstrate that ANESS significantly outperforms other state-of-the-art network embedding methods.
Read moreAttributed Network Embedding with Micro-Meso Structure
Recently, network embedding has received a large amount of attention in network analysis. Although some network embedding methods have been developed from different perspectives, on one hand, most of the existing methods only focus on leveraging the plain network structure, ignoring the abundant attribute information of nodes. On the other hand, for some methods integrating the attribute information, only the lower-order proximities (e.g., microscopic proximity structure) are taken into account, which may suffer if there exists the sparsity issue and the attribute information is noisy. To overcome this problem, the attribute information and mesoscopic community structure are utilized. In this article, we propose a novel network embedding method termed Attributed Network Embedding with Micro-Meso structure, which is capable of preserving both the attribute information and the structural information including the microscopic proximity structure and mesoscopic community structure. In particular, both the microscopic proximity structure and node attributes are factorized by Nonnegative Matrix Factorization (NMF), from which the low-dimensional node representations can be obtained. For the mesoscopic community structure, a community membership strength matrix is inferred by a generative model (i.e., BigCLAM) or modularity from the linkage structure, which is then factorized by NMF to obtain the low-dimensional node representations. The three components are jointly correlated by the low-dimensional node representations, from which two objective functions (i.e., ANEM_B and ANEM_M) can be defined. Two efficient alternating optimization schemes are proposed to solve the optimization problems. Extensive experiments have been conducted to confirm the superior performance of the proposed models over the state-of-the-art network embedding methods.
Read moreNetwork-Word Embedding for Dynamic Text Attributed Networks
Network embedding enables to apply off-the-shelf machine learning methods to the nodes on the network. Leveraging the textual information associated with nodes into network embedding methods is advantageous. However, only a few works try to leverage textual information to network embeddings. Moreover, the structure of networks and associated texts could be dynamically changed over time, this property leads shifts of the vector representation of nodes and words. However, to the best of our knowledge, none of previous network embedding methods considers chronological changes of vector representations of nodes and words. In this paper, we propose a dynamic text attribute network embedding method, which embeds the nodes and the words in a cooperative manner and takes chronological changes of vector representations of nodes and words into account. Experimental results show that (1) vector representations of nodes of our method achieve higher accuracy than baseline methods in classification and clustering tasks; (2) The vector representations of nodes and words successfully capture semantic similarity; And (3) our method successfully capture the chronological change of the vector representations over time.
Read moreSINE: Second-Order Information Network Embedding
As an important data type, the demand of network analysis and learning is increasingly prominent. A key problem of network analysis is to study how to reasonably represent the feature information in the network, that is network embedding. However, in the study of network embedding, only the first-order proximity relationship of nodes is characterized, while the second-order proximity relationship hidden in the network nodes is ignored. Therefore, we propose an algorithm, called SINE, to realize the representation learning of nodes in the network which fuses the first-order and second-order proximity of nodes in the original network. Through applying SINE to three real networks, we obtain the feature representation of the nodes in the networks and cluster the nodes based on these features. Comparing with the existing network embedding learning methods-Node2vec, Large-Scale Information Network Embedding (LINE) and Structural Deep Network Embedding (SDNE), the SINE algorithm showed better performance in clustering tasks.
Read moreCompact network embedding for fast node classification
Compact network embedding for fast node classification
Community Preserving Network Embedding
Network embedding, aiming to learn the low-dimensional representations of nodes in networks, is of paramount importance in many real applications. One basic requirement of network embedding is to preserve the structure and inherent properties of the networks. While previous network embedding methods primarily preserve the microscopic structure, such as the first- and second-order proximities of nodes, the mesoscopic community structure, which is one of the most prominent feature of networks, is largely ignored. In this paper, we propose a novel Modularized Nonnegative Matrix Factorization (M-NMF) model to incorporate the community structure into network embedding. We exploit the consensus relationship between the representations of nodes and community structure, and then jointly optimize NMF based representation learning model and modularity based community detection model in a unified framework, which enables the learned representations of nodes to preserve both of the microscopic and community structures. We also provide efficient updating rules to infer the parameters of our model, together with the correctness and convergence guarantees. Extensive experimental results on a variety of real-world networks show the superior performance of the proposed method over the state-of-the-arts.
Read moreAdaptive Attributed Network Embedding for Community Detection
Community detection, which discovers densely-connected groups of nodes in networks, is a fundamental task in machine learning and data mining. Compared with plain network, community detection in attributed network presents more challenges. Several recent embedding-based methods have achieved promising community detection performance on some real attributed networks. However, there is limited understanding of how to effectively learn the combination of heterogeneity in the joint space of topology and attribute in unsupervised scenarios. In this paper, we propose an end-to-end network embedding method. By employing the high order graph convolutional networks, our method encodes the topological structure and node attributes to learn compact representations, i.e., community membership. The decoder on the other side, we reconstruct the global topological structure based on the learned community membership and stochastic block model. We further employ a self-training module, which takes the “confident” link assignments as soft labels to guide the optimizing procedure. Experiments show our method has achieved the sate-of-the-art performance on three popular datasets.
Read moreDiscrete Embedding for Latent Networks
Discrete network embedding emerged recently as a new direction of network representation learning. Compared with traditional network embedding models, discrete network embedding aims to compress model size and accelerate model inference by learning a set of short binary codes for network vertices. However, existing discrete network embedding methods usually assume that the network structures (e.g., edge weights) are readily available. In real-world scenarios such as social networks, sometimes it is impossible to collect explicit network structure information and it usually needs to be inferred from implicit data such as information cascades in the networks. To address this issue, we present an end-to-end discrete network embedding model for latent networks DELN that can learn binary representations from underlying information cascades. The essential idea is to infer a latent Weisfeiler-Lehman proximity matrix that captures node dependence based on information cascades and then to factorize the latent Weisfiler-Lehman matrix under the binary node representation constraint. Since the learning problem is a mixed integer optimization problem, an efficient maximal likelihood estimation based cyclic coordinate descent (MLE-CCD) algorithm is used as the solution. Experiments on real-world datasets show that the proposed model outperforms the state-of-the-art network embedding methods.
Read moreOn Network Embedding for Machine Learning on Road Networks: A Case Study on the Danish Road Network
Road networks are a type of spatial network, where edges may be associated with qualitative information such as road type and speed limit. Unfortunately, such information is often incomplete; for instance, OpenStreetMap only has speed limits for 13% of all Danish road segments. This is problematic for analysis tasks that rely on such information for machine learning. To enable machine learning in such circumstances, one may consider the application of network embedding methods to extract structural information from the network. However, these methods have so far mostly been used in the context of social networks, which differ significantly from road networks in terms of, e.g., node degree and level of homophily (which are key to the performance of many network embedding methods). We analyze the use of network embedding methods, specifically node2vec, for learning road segment embeddings in road networks. Due to the often limited availability of information on other relevant road characteristics, the analysis focuses on leveraging the spatial network structure. Our results suggest that network embedding methods can indeed be used for deriving relevant network features (that may, e.g, be used for predicting speed limits), but that the qualities of the embeddings differ from embeddings for social networks.
Read moreNetwork Embedding via Link Strength Adjusted Random Walk
Network embedding is a useful tool to map graph structures into vector spaces, which facilitates graph analysis tasks including node classification, graph visualization, similarity calculation etc. Existing network embedding methods calculate embedding vectors based on node series generated by random walks. These methods treat all the links equally during the random walk procedure, which leads to the missing of structural information that is key to the embedding performance. We therefore propose in this paper a novel random walk-based network embedding method called Self-Adjusting Random Walk (SARW). SARW utilizes a self-adjusting strategy that makes the walking biased towards the links that are more strongly connected in order to better capture the structural information. Further more, the strengths of links are updated using the embedding output as feedback. Through experiments we have verified that our method out performs state-of-the-art network embedding methods, in node classification tasks and link prediction tasks.
Read moreDeep Attributed Network Embedding by Preserving Structure and Attribute Information
Network embedding aims to learn distributed vector representations of nodes in a network. The problem of network embedding is fundamentally important. It plays crucial roles in many applications, such as node classification, link prediction, and so on. As the real-world networks are often sparse with few observed links, many recent works have utilized the local and global network structure proximity with shallow models for better network embedding. In reality, each node is usually associated with rich attributes. Some attributed network embedding models leveraged the node attributes in these shallow network embedding models to alleviate the data sparsity issue. Nevertheless, the underlying structure of the network is complex. What is more, the connection between the network structure and node attributes is also hidden. Thus, these previous shallow models fail to capture the nonlinear deep information embedded in the attributed network, resulting in the suboptimal embedding results. In this paper, we propose a deep attributed network embedding framework to capture the complex structure and attribute information. Specifically, we first adopt a personalized random walk-based model to capture the interaction between network structure and node attributes from various degrees of proximity. After that, we construct an enhanced matrix representation of the attributed network by summarizing the various degrees of proximity. Then, we design a deep neural network to exploit the nonlinear complex information in the enhanced matrix for network embedding. Thus, the proposed framework could capture the complex attributed network structure by preserving both the various degrees of network structure and node attributes in a unified framework. Finally, empirical experiments show the effectiveness of our proposed framework on a variety of network embedding-based tasks.
Read moreRNE: A Scalable Network Embedding for Billion-Scale Recommendation
Nowadays designing a real recommendation system has been a critical problem for both academic and industry. However, due to the huge number of users and items, the diversity and dynamic property of the user interest, how to design a scalable recommendation system, which is able to efficiently produce effective and diverse recommendation results on billion-scale scenarios, is still a challenging and open problem for existing methods. In this paper, given the user-item interaction graph, we propose RNE, a data-efficient Recommendation-based Network Embedding method, to give personalized and diverse items to users. Specifically, we propose a diversity- and dynamics-aware neighbor sampling method for network embedding. On the one hand, the method is able to preserve the local structure between the users and items while modeling the diversity and dynamic property of the user interest to boost the recommendation quality. On the other hand the sampling method can reduce the complexity of the whole method theoretically to make it possible for billion-scale recommendation. We also implement the designed algorithm in a distributed way to further improves its scalability. Experimentally, we deploy RNE on a recommendation scenario of Taobao, the largest E-commerce platform in China, and train it on a billion-scale user-item graph. As is shown on several online metrics on A/B testing, RNE is able to achieve both high-quality and diverse results compared with CF-based methods. We also conduct the offline experiments on Pinterest dataset comparing with several state-of-the-art recommendation methods and network embedding methods. The results demonstrate that our method is able to produce a good result while runs much faster than the baseline methods.
Read more