- Research Article
7
- 10.1016/j.neucom.2024.127523
Deep contrastive representation learning for multi-modal clustering
- Mar 06, 2024
- Neurocomputing
- Yang Lu + 3 more +3
Deep contrastive representation learning for multi-modal clustering
Multi-modal clustering (MMC) aims to explore complementary information from diverse modalities for clustering performance facilitating. This article studies challenging problems in MMC methods based on deep neural networks. On one hand, most existing methods lack a unified objective to simultaneously learn the inter- and intra-modality consistency, resulting in a limited representation learning capacity. On the other hand, most existing processes are modeled for a finite sample set and cannot handle out-of-sample data. To handle the above two challenges, we propose a novel Graph Embedding Contrastive Multi-modal Clustering network (GECMC), which treats the representation learning and multi-modal clustering as two sides of one coin rather than two separate problems. In brief, we specifically design a contrastive loss by benefiting from pseudo-labels to explore consistency across modalities. Thus, GECMC shows an effective way to maximize the similarities of intra-cluster representations while minimizing the similarities of inter-cluster representations at both inter- and intra-modality levels. So, the clustering and representation learning interact and jointly evolve in a co-training framework. After that, we build a clustering layer parameterized with cluster centroids, showing that GECMC can learn the clustering labels with given samples and handle out-of-sample data. GECMC yields superior results than 14 competitive methods on four challenging datasets. Codes and datasets are available: https://github.com/xdweixia/GECMC.
Deep contrastive representation learning for multi-modal clustering
Deep contrastive representation learning for multi-modal clustering
Multimodal and Crossmodal Representation Learning from Textual and Visual Features with Bidirectional Deep Neural Networks for Video Hyperlinking
Video hyperlinking represents a classical example of multimodal problems. Common approaches to such problems are early fusion of the initial modalities and crossmodal translation from one modality to the other. Recently, deep neural networks, especially deep autoencoders, have proven promising both for crossmodal translation and for early fusion via multimodal embedding. A particular architecture, bidirectional symmetrical deep neural networks, have been proven to yield improved multimodal embeddings over classical autoencoders, while also being able to perform crossmodal translation. In this work, we focus firstly at evaluating good single-modal continuous representations both for textual and for visual information. Word2Vec and paragraph vectors are evaluated for representing collections of words, such as parts of automatic transcripts and multiple visual concepts, while different deep convolutional neural networks are evaluated for directly embedding visual information, avoiding the creation of visual concepts. Secondly, we evaluate methods for multimodal fusion and crossmodal translation, with different single-modal pairs, in the task of video hyperlinking. Bidirectional (symmetrical) deep neural networks were shown to successfully tackle downsides of multimodal autoencoders and yield a superior multimodal representation. In this work, we extensively tests them in different settings, with different single-modal representations, within the context of video-hyperlinking. Our novel bidirectional symmetrical deep neural networks are compared to classical autoencoders and are shown to yield significantly improved multimodal embeddings that significantly (alpha=0.0001) outperform multimodal embeddings obtained by deep autoencoders with an absolute improvement in precision at 10 of 14.1% when embedding visual concepts and automatic transcripts and an absolute improvement of 4.3% when embedding automatic transcripts with features obtained with very deep convolutional neural networks, yielding 80% of precision at 10.
Read moreMulti-modal knowledge graphs representation learning via multi-headed self-attention
Multi-modal knowledge graphs representation learning via multi-headed self-attention
Multimodal Representation Learning via Graph Isomorphism Network for Toxicity Multitask Learning.
Toxicity is paramount for comprehending compound properties, particularly in the early stages of drug design. Due to the diversity and complexity of toxic effects, it became a challenge to compute compound toxicity tasks. To address this issue, we propose a multimodal representation learning model, termed multimodal graph isomorphism network (MMGIN), to address this challenge for compound toxicity multitask learning. Based on fingerprints and molecular graphs of compounds, our MMGIN model incorporates a multimodal representation learning model to acquire a comprehensive compound representation. This model adopts a two-channel structure to independently learn fingerprint representation and molecular graph representation. Subsequently, two feedforward neural networks utilize the learned multimodal compound representation to perform multitask learning, encompassing compound toxicity classification and multiple compound category classification simultaneously. To test the effectiveness of our model, we constructed a novel data set, termed the compound toxicity multitask learning (CTMTL) data set, derived from the TOXRIC data set. We compare our MMGIN model with other representative machine learning and deep learning models on the CTMTL and Tox21 data sets. The experimental results demonstrate significant advancements achieved by our MMGIN model. Furthermore, the ablation study underscores the effectiveness of the introduced fingerprints, molecular graphs, the multimodal representation learning model, and the multitask learning model, showcasing the model's superior predictive capability and robustness.
Read moreDeep soft clustering: simultaneous deep embedding and soft-partition clustering
Traditional clustering methods are not very effective when dealing with high-dimensional and huge datasets. Even if there are some traditional dimensionality reduction methods such as Principal components analysis (PCA), Linear discriminant analysis (LDA) and T-distributed stochastic neighbor embedding (T-SNE), they still can not significantly improve the effect of the clustering algorithm in this scenario. Recent studies have combined Non-linear dimensionality reduction achieved by deep neural networks with hard-partition clustering, and have achieved reliability results, but these methods can not update the parameters of dimensionality reduction and clustering at the same time. We found that soft-partition clustering can be well combined with deep embedding, and the membership of Fuzzy c-means (FCM) can solve the problem that gradient descent can not be implemented because the assignment process of the hard-partition clustering algorithm is discrete, so that the algorithm can update the parameters of deep neural network (DNN) and cluster centroids at the same time. We build an continuous objective function that combine the soft-partition clustering with deep embedding, so that the learning representations can be cluster-friendly. The experimental results show that our proposed method of simultaneously optimizing the parameters of deep dimensionality reduction and clustering is better than the method with separate optimization.
Read moreMultimodal deep learning for solar radio burst classification
Multimodal deep learning for solar radio burst classification
Representation learning for geospatial data
This paper reviews representation learning for geospatial data, focusing on methods for automatically extracting meaningful features from diverse data types. By simplifying tasks and improving accuracy, representation learning has emerged as a powerful tool for geospatial analysis. Due to its generalizability and scalability, representation learning provides an effective approach to processing geospatial data, which is inherently diverse and unstructured. We summarize the representation learning methods for different geospatial data types, including locations, points of interest (POIs), trajectories, spatial interactions, remote sensing imagery, and street view imagery. Treating each data type as a distinct modality, we emphasize the potential of multi-modal representation learning to advance the understanding of geographical phenomena and propose an LLM-guided framework as a potential solution. The review concludes by highlighting the need for further research to improve multi-modal data alignment and enhance the interpretability of feature representations, particularly in complex and dynamic geographical environments.
Read moreHuman-centric 3D representation learning
Understanding the world in three dimensions has long been a scientific challenge with significant practical implications. In the era of deep learning, researchers have made substantial progress in developing 3D representations using deep neural networks. Among the myriad of entities in our complex environment, humans stand out due to their general significance in science and their specific relevance in various applications. This thesis focuses on human-centric 3D representation learning within the framework of deep learning. Specifically, we explore three key areas: human-centric 3D perception, human reconstruction, and 3D human generation. The thesis presents five studies that collectively address these topics. In the realm of human-centric 3D perception, we introduce a versatile multi-modal pre-training approach. By harnessing the diverse modalities of human data, such as RGB images, depth, and 2D keypoints, we present a general framework HCMoCo for effective human-centric representation learning. This framework achieves state-of-the-art performance across four human perception tasks, including DensePose prediction, human parsing, and 3D keypoint prediction. In the area of human reconstruction, we investigate garment reconstruction from 4D point clouds of dressed individuals. To address the ambiguities inherent in 2D images, we propose a principled framework called Garment4D, which enables separable and interpretable garment reconstruction. Notably, Garment4D can effectively reconstruct and model the non-rigid deformations of loose garments (e.g., skirts) that do not share the same topology as the human body. We present two distinct approaches to 3D human generation. In our first work, AvatarCLIP, we generate and animate avatars based on text descriptions of body shapes, appearances, and motions. By leveraging differentiable rendering and large- scale vision-language pre-trained models, we achieve avatar and motion synthesis without the need for supervised training or paired data. To further enhance the quality of the generated avatars, we introduce a second approach, EVA3D, which is a high-quality unconditional 3D human generative model that requires only 2D image collections for training. We design an efficient compositional human NeRF representation to facilitate high-resolution 3D human sampling and rendering, which is employed in adversarial training. Finally, by harnessing the capabilities of large language models (LLMs), we explore 3D human representation learning with a focus on human motion and the integration of perception, reconstruction, and generation. We introduce EgoLM, a versatile framework for understanding egocentric motion using multi-modal data. This framework incorporates rich contextual information from egocentric videos and motion sensors provided by wearable devices. EgoLM unifies various motion learning tasks, including motion understanding from video and motion data, as well as motion tracking and generation from text or sparse sensor input. Additionally, it enables a novel task unique to wearable devices: generating text descriptions from sparse sensor data. By investigating 3D human representation learning from three distinct perspectives, we have attained a comprehensive understanding of humans that spans both high-level semantics and low-level geometry. This thesis offers a cohesive exploration of human-centric 3D representations, contributing significantly to the field of human-centric vision.
Read moreMultimodal Intelligence: Representation Learning, Information Fusion, and Applications
Deep learning methods have revolutionized speech recognition, image recognition, and natural language processing since 2010. Each of these tasks involves a single modality in their input signals. However, many applications in the artificial intelligence field involve multiple modalities. Therefore, it is of broad interest to study the more difficult and complex problem of modeling and learning across multiple modalities. In this paper, we provide a technical review of available models and learning methods for multimodal intelligence. The main focus of this review is the combination of vision and natural language modalities, which has become an important topic in both the computer vision and natural language processing research communities. This review provides a comprehensive analysis of recent works on multimodal deep learning from three perspectives: learning multimodal representations, fusing multimodal signals at various levels, and multimodal applications. Regarding multimodal representation learning, we review the key concepts of embedding, which unify multimodal signals into a single vector space and thereby enable cross-modality signal processing. We also review the properties of many types of embeddings that are constructed and learned for general downstream tasks. Regarding multimodal fusion, this review focuses on special architectures for the integration of representations of unimodal signals for a particular task. Regarding applications, selected areas of a broad interest in the current literature are covered, including image-to-text caption generation, text-to-image generation, and visual question answering. We believe that this review will facilitate future studies in the emerging field of multimodal intelligence for related communities.
Read moreDual-Stage Clean-Sample Selection for Incremental Noisy Label Learning.
Class-incremental learning (CIL) in deep neural networks is affected by catastrophic forgetting (CF), where acquiring knowledge of new classes leads to the significant degradation of previously learned representations. This challenge is particularly severe in medical image analysis, where costly, expertise-dependent annotations frequently contain pervasive and hard-to-detect noisy labels that substantially compromise model performance. While existing approaches have predominantly addressed CF and noisy labels as separate problems, their combined effects remain largely unexplored. To address this critical gap, this paper presents a dual-stage clean-sample selection method for Incremental Noisy Label Learning (DSCNL). Our approach comprises two key components: (1) a dual-stage clean-sample selection module that identifies and leverages high-confidence samples to guide the learning of reliable representations while mitigating noise propagation during training, and (2) an experience soft-replay strategy for memory rehearsal to improve the model's robustness and generalization in the presence of historical noisy labels. This integrated framework effectively suppresses the adverse influence of noisy labels while simultaneously alleviating catastrophic forgetting. Extensive evaluations on public medical image datasets demonstrate that DSCNL consistently outperforms state-of-the-art CIL methods across diverse classification tasks. The proposed method boosts the average accuracy by 55% and 31% compared with baseline methods on datasets with different noise levels, and achieves an average noise reduction rate of 73% under original noise conditions, highlighting its effectiveness and applicability in real-world medical imaging scenarios.
Read moreMAMO: Fine-Grained Vision-Language Representations Learning with Masked Multimodal Modeling
Multimodal representation learning has shown promising improvements on various vision-language tasks (e.g., image-text retrieval, visual question answering, etc) and has significantly advanced the development of multimedia information systems. Most existing methods excel at building global-level alignment between vision and language while lacking effective fine-grained image-text interaction. In this paper, we propose a jointly masked multimodal modeling method to learn fine-grained multimodal representations. Our method performs joint masking on image-text input and integrates both implicit and explicit targets for the masked signals to recover. The implicit target provides a unified and debiased objective for vision and language, where the model predicts latent multimodal representations of the unmasked input. The explicit target further enriches the multimodal representations by recovering high-level and semantically meaningful information: momentum visual features of image patches and concepts of word tokens. Through such a masked modeling process, our model not only learns fine-grained multimodal interaction, but also avoids the semantic gap between high-level representations and low-or mid-level prediction targets (e.g., image pixels, discrete vision tokens), thus producing semantically rich multimodal representations that perform well on both zero-shot and fine-tuned settings. Our pre-trained model (named MAMO) achieves state-of-the-art performance on various downstream vision-language tasks, including image-text retrieval, visual question answering, visual reasoning, and weakly-supervised visual grounding.
Read moreLearning Multimodal Representations by Symmetrically Transferring Local Structures
Multimodal representations play an important role in multimodal learning tasks, including cross-modal retrieval and intra-modal clustering. However, existing multimodal representation learning approaches focus on building one common space by aligning different modalities and ignore the complementary information across the modalities, such as the intra-modal local structures. In other words, they only focus on the object-level alignment and ignore structure-level alignment. To tackle the problem, we propose a novel symmetric multimodal representation learning framework by transferring local structures across different modalities, namely MTLS. A customized soft metric learning strategy and an iterative parameter learning process are designed to symmetrically transfer local structures and enhance the cluster structures in intra-modal representations. The bidirectional retrieval loss based on multi-layer neural networks is utilized to align two modalities. MTLS is instantiated with image and text data and shows its superior performance on image-text retrieval and image clustering. MTLS outperforms the state-of-the-art multimodal learning methods by up to 32% in terms of R@1 on text-image retrieval and 16.4% in terms of AMI onclustering.
Read moreAn Optimal Transport-based Latent Mixer for Robust Multi-modal Learning
Multi-modal learning aims to learn predictive models based on the data from different modalities. However, due to the requirement of data security and privacy protection, real-world multi-modal data are often scattered to different agents and cannot be shared across the agents, which limits the application of existing multi-modal learning methods. To achieve robust multi-modal learning in such a challenging scenario, we propose a novel optimal transport-based mixer (OTM), which works as an effective latent code alignment and augmentation method for unaligned and distributed multi-modal data. In particular, we train a Wasserstein autoencoder (WAE) for each agent, which encodes its single modal samples in a latent space. Through a central server, the proposed OTM computes a stochastic fused Gromov-Wasserstein barycenter (FGWB) to mix different modalities' latent codes, so that each agent applies the barycenter to reconstruct its samples. This method neither requires well-aligned multi-modal data nor assumes the data to share the same latent distribution, and each agent can learn a specific model based on multi-modal data while achieving inference based on its local modality. Experiments on multi-modal clustering and classification demonstrate that the models learned with the OTM method outperform the corresponding baselines.
Read moreLearning Eligibility in Cancer Clinical Trials Using Deep Neural Networks
Interventional cancer clinical trials are generally too restrictive, and some patients are often excluded on the basis of comorbidity, past or concomitant treatments, or the fact that they are over a certain age. The efficacy and safety of new treatments for patients with these characteristics are, therefore, not defined. In this work, we built a model to automatically predict whether short clinical statements were considered inclusion or exclusion criteria. We used protocols from cancer clinical trials that were available in public registries from the last 18 years to train word-embeddings, and we constructed a dataset of 6M short free-texts labeled as eligible or not eligible. A text classifier was trained using deep neural networks, with pre-trained word-embeddings as inputs, to predict whether or not short free-text statements describing clinical information were considered eligible. We additionally analyzed the semantic reasoning of the word-embedding representations obtained and were able to identify equivalent treatments for a type of tumor analogous with the drugs used to treat other tumors. We show that representation learning using deep neural networks can be successfully leveraged to extract the medical knowledge from clinical trial protocols for potentially assisting practitioners when prescribing treatments.
Read moreAn Empirical Study on the Countermeasures of Implementing 5G Multimedia Network Technology in College Education
Aiming at the problem of 5G multimedia heterogeneous multimodal network representation learning, this paper proposes a collaborative multimodal heterogeneous network representation learning method based on attention mechanism. This method learns different representations for nodes based on heterogeneous network structure information and multimodal content and designs an attention mechanism to learn weights for different representations to fuse them to obtain robust node representations. Combining the general process of exploring the college physical education model and the characteristics of the multimedia network classroom environment, this article constructs the process of exploring the college physical education teaching model of the multimedia network classroom. Through the research and practice of the inquiry college physical education teaching model in the multimedia network classroom, it is verified that the implementation of the inquiry college physical education teaching in the multimedia network classroom can effectively influence and increase the students’ interest in learning and stimulate the students’ inner learning motivation. Through the guidance and training of teachers, a variety of disciplines can be used to carry out college physical education in multimedia network classrooms, so that the integration between courses can be truly realized, with the aim that all courses can share the excellent results brought by the development of modern education technology. More educators understand, accept, and participate in the practice of college physical education based on multimedia network classrooms and better serve the education of college physical education. The construction of the college physical education evaluation system should be combined with the characteristics of the 5G multimedia network era. The evaluation process includes data collection, data analysis, result output, and result feedback. Each link is an indispensable part of the college physical education evaluation process. Based on the relevant knowledge of the 5G multimedia network, the evaluation indicators determined in this study can basically reflect the various elements of the physical education process in colleges and universities. The distribution of index weight coefficients is more scientific and reasonable. Compared with the current system, the college physical education evaluation system constructed by exploration has a certain degree of objectivity and scientificity. Therefore, it is feasible to apply the 5G multimedia network to the evaluation of college physical education.
Read more