- Research Article
13
- 10.1016/j.knosys.2023.110899
Semantically consistent multi-view representation learning
- Aug 11, 2023
- Knowledge-Based Systems
- Yiyang Zhou + 3 more +3
Semantically consistent multi-view representation learning
Multiview data, characterized by rich features, are crucial in many machine learning applications. However, effectively extracting intraview features and integrating interview information present significant challenges in multiview learning (MVL). Traditional deep network-based approaches often involve learning multiple layers to derive latent. In these methods, the features of different classes are typically implicitly embedded rather than systematically organized. This lack of structure makes it challenging to explicitly map classes to independent principal subspaces in the feature space, potentially causing class overlap and confusion. Consequently, the capability of these representations to accurately capture the intrinsic structure of the data remains uncertain. In this article, we introduce an innovative multiview representation learning (MVRL) by maximizing two information-theoretic metrics: intraview coding rate reduction and interview mutual information. Specifically, in the intraview representation learning, we aim to optimize feature representations by maximizing the coding rate difference between the entire dataset and individual classes. This process expands the feature representation space while compressing the representations within each class, resulting in more compact feature representations within each viewpoint. Subsequently, we align and fuse these view-specific features through space transformation and cross-sample fusion to achieve consistent representation across multiple views. Finally, we maximize information transmission to maintain consistency and correlation among data representations across views. By maximizing mutual information between the consensus representations and view-specific representations, our method ensures that the learned representations capture more concise intrinsic features and correlations among different views, thereby enhancing the performance and generalization ability of MVL. Experiments show that the proposed methods have achieved excellent performance.
Semantically consistent multi-view representation learning
Semantically consistent multi-view representation learning
Self-Supervised Deep Correlational Multi-View Clustering
In conventional unsupervised multi-view clustering (MVC), learning of representations from heterogeneous multiview data and its subsequent clustering are often separately optimized. The disparate optimization would lead to suboptimal performance because multi-view representation learning is not goal-directed. In this paper, we unify unsupervised multi-view learning and deep clustering in a novel discriminative Self-supervised Deep Correlational Multi-view Clustering (SDC-MVC) network. A new unified loss function is proposed to incorporate consensus information into discriminative representations, in which, the former is learnt by maximizing the canonical correlation among multi-view representations projected by neural networks, and the later is achieved through using confident clustering assignments as supervision. Further, multi-view representations are harnessed by our proposed Deep Serial Feature-level (DSF) Fusion layer. Experiments on three public datasets demonstrated that our method outperforms six state-of-the-art correlation-based MVC algorithms in terms of three evaluation metrics.
Read moreVideo semantic segmentation using deep multi-view representation learning
In this paper, we propose a deep learning model based on deep multi-view representation learning, to address the video object segmentation task. The proposed model emphasizes the importance of the inherent correlation between video frames and incorporates a multi-view representation learning based on deep canonically correlated autoencoders. The multi-view representation learning in our model provides an efficient mechanism for capturing inherent correlations by jointly extracting useful features and learning better representation into a joint feature space, i.e., shared representation. To increase the training data and the learning capacity, we train the proposed model with pairs of video frames, i.e., Fa and Fb. During the segmentation phase, the deep canonically correlated autoencoders model encodes useful features by processing multiple reference frames together, which is used to detect the frequently reappearing. Our model enhances the state-of-the-art deep learning-based methods that mainly focus on learning discriminative foreground representations over appearance and motion. Experimental results over two large benchmarks demonstrate the ability of the proposed method to outperform competitive approaches and to reach good performances, in terms of semantic segmentation.
Read morePoint Context: An Effective Shape Descriptor for RST-Invariant Trajectory Recognition
Motion trajectory recognition is important for characterizing the moving property of an object. The speed and accuracy of trajectory recognition rely on a compact and discriminative feature representation, and the situations of varying rotation, scaling and translation has to be specially considered. In this paper we propose a novel feature extraction method for trajectories. Firstly a trajectory is represented by a proposed point context, which is a rotation-scale-translation (RST) invariant shape descriptor with a flexible tradeoff between computational complexity and discrimination, yet we prove that it is a complete shape descriptor. Secondly, the shape context is nonlinearly mapped to a subspace by kernel nonparametric discriminant analysis (KNDA) to get a compact feature representation, and thus a trajectory is projected to a single point in a low-dimensional feature space. Experimental results show that, the proposed trajectory feature shows encouraging improvement than state-of-art methods.
Read moreAdversarial correlated autoencoder for unsupervised multi-view representation learning
Adversarial correlated autoencoder for unsupervised multi-view representation learning
Multi-View Information-Bottleneck Representation Learning
In real-world applications, clustering or classification can usually be improved by fusing information from different views. Therefore, unsupervised representation learning on multi-view data becomes a compelling topic in machine learning. In this paper, we propose a novel and flexible unsupervised multi-view representation learning model termed Collaborative Multi-View Information Bottleneck Networks (CMIB-Nets), which comprehensively explores the common latent structure and the view-specific intrinsic information, and discards the superfluous information in the data significantly improving the generalization capability of the model. Specifically, our proposed model relies on the information bottleneck principle to integrate the shared representation among different views and the view-specific representation of each view, prompting the multi-view complete representation and flexibly balancing the complementarity and consistency among multiple views. We conduct extensive experiments (including clustering analysis, robustness experiment, and ablation study) on real-world datasets, which empirically show promising generalization ability and robustness compared to state-of-the-arts.
Read moreMulti-view subspace learning via bidirectional sparsity
Multi-view subspace learning via bidirectional sparsity
Tensor Discriminant Analysis via Compact Feature Representation for Hyperspectral Images Dimensionality Reduction
Dimensionality reduction is of great importance which aims at reducing the spectral dimensionality while keeping the desirable intrinsic structure information of hyperspectral images. Tensor analysis which can retain both spatial and spectral information of hyperspectral images has caused more and more concern in the field of hyperspectral images processing. In general, a desirable low dimensionality feature representation should be discriminative and compact. To achieve this, a tensor discriminant analysis model via compact feature representation (TDA-CFR) was proposed in this paper. In TDA-CFR, the traditional linear discriminant analysis was extended to tensor space to make the resulting feature representation more informative and discriminative. Furthermore, TDA-CFR redefines the feature representation of each spectral band by employing the tensor low rank decomposition framework which leads to a more compact representation.
Read moreMulti-view representation learning with Kolmogorov-Smirnov to predict default based on imbalanced and complex dataset
Multi-view representation learning with Kolmogorov-Smirnov to predict default based on imbalanced and complex dataset
Multi-view representation learning for performance-based tactical conflict resolution in urban air mobility
Ensuring safety and security in urban air mobility is of utmost significance. As air traffic becomes more concentrated in urban regions, instances of flight conflicts are on the rise. The complex urban morphology, micro-environmental factors and various flight risks significantly impact flight safety. This paper introduces a comprehensive framework to address these challenges. The framework incorporates random 3D city layouts, encompassing city buildings and terrains to establish a corridor graph structure. Leveraging this graph, a multi-view representation learning approach is proposed, which employs graph neural networks, recurrent neural networks (RNN) and contrastive learning to effectively manage tactical conflicts. Through rigorous testing across diverse scenarios, including the incorporation of uncertainties such as wind turbulence, the model’s performance is extensively evaluated. The conclusive results underscore the robustness and efficacy of the proposed approach in ensuring safety and resolving conflicts within the dynamic urban air mobility landscape.
Read moreMulti-view Low-rank Preserving Embedding: A novel method for multi-view representation
Multi-view Low-rank Preserving Embedding: A novel method for multi-view representation
Kernelized multi-view subspace clustering via auto-weighted graph learning
Multi-view subspace clustering has been an important and powerful tool for partitioning multi-view data, especially multi-view high-dimensional data. Despite great success, most of the existing multi-view subspace clustering methods still suffer from three limitations. First, they often recover the subspace structure in the original space, which can not guarantee the robustness when handling multi-view data with nonlinear structure. Second, these methods mostly regard subspace clustering and affinity matrix learning as two independent steps, which may not well discover the latent relationships among data samples. Third, many of them ignore the different importance of multiple views, whose performance may be badly affected by the low-quality views in multi-view data. To overcome these three limitations, this paper develops a novel subspace clustering method for multi-view data, termed Kernelized Multi-view Subspace Clustering via Auto-weighted Graph Learning (KMSC-AGL). Specifically, the proposed method implicitly maps the multi-view data from linear space into nonlinear space via kernel-induced functions, so as to exploit the nonlinear structure hidden in data. Furthermore, our method aims to enhance the clustering performance by learning a set of view-specific representations and their affinity matrix in a general framework. By integrating the view weighting strategy into this framework, our method can automatically assign the weights to different views, while learning an optimal affinity matrix that is well-adapted to the subsequent spectral clustering. Extensive experiments are conducted on a variety of multi-view data sets, which have demonstrated the superiority of the proposed method.
Read morePH-GCN: Person Retrieval With Part-Based Hierarchical Graph Convolutional Network
Compact feature representation of person image is important for person re-identification (Re-ID) task. Recently, part-based representation models have been widely studied for extracting the more compact and robust feature representation for person image to improve person Re-ID results. However, existing part-based representation models mostly extract the features of different parts independently which ignore the spatial relationship information among different parts. To address this issue, in this paper we propose a novel deep learning framework, named Part-based Hierarchical Graph Convolutional Network (PH-GCN) for person Re-ID problem. Given a person image, PH-GCN first constructs a hierarchical graph to represent the spatial relationships among different parts. Then, both local and global feature learning is achieved by the feature information passing in PH-GCN, which takes the information of other parts into account for part feature representation. Finally, a perceptron layer is adopted for the final person part label prediction and re-identification. The proposed framework provides a general solution that integrates <i>local</i>, <i>global</i> and <i>structural</i> feature learning simultaneously in a unified end-to-end network representation and learning. Extensive experiments on several widely used benchmark datasets demonstrate the effectiveness and benefits of the proposed PH-GCN approach for person Re-ID task.
Read moreUsing unlabeled data in a sparse-coding framework for human activity recognition
Using unlabeled data in a sparse-coding framework for human activity recognition
Multi-View Cosine Similarity Learning with Application to Face Verification
An instance can be easily depicted from different views in pattern recognition, and it is desirable to exploit the information of these views to complement each other. However, most of the metric learning or similarity learning methods are developed for single-view feature representation over the past two decades, which is not suitable for dealing with multi-view data directly. In this paper, we propose a multi-view cosine similarity learning (MVCSL) approach to efficiently utilize multi-view data and apply it for face verification. The proposed MVCSL method is able to leverage both the common information of multi-view data and the private information of each view, which jointly learns a cosine similarity for each view in the transformed subspace and integrates the cosine similarities of all the views in a unified framework. Specifically, MVCSL employs the constraints that the joint cosine similarity of positive pairs is greater than that of negative pairs. Experiments on fine-grained face verification and kinship verification tasks demonstrate the superiority of our MVCSL approach.
Read more