• Home
  • Search
  • Vision-Language Pre-Training with Triple Contrastive Learning
  • Open Access IconOpen Access
  • Cite Icon265
  • https://doi.org/10.1109/cvpr52688.2022.01522Copy DOI Icon

Vision-Language Pre-Training with Triple Contrastive Learning

  • Jun 1, 2022
  • Jinyu Yang +8 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attributed to its capability in maximizing the mutual information (MI) between an image and its matched text. However, simply performing cross-modal alignment (CMA) ignores data potential within each modality, which may result in degraded representations. For instance, although CMA-based models are able to map image-text pairs close together in the embedding space, they fail to ensure that similar inputs from the same modality stay close by. This problem can get even worse when the pre-training data is noisy. In this paper, we propose triple contrastive learning (TCL) for vision-language pre-training by leveraging both cross-modal and intra-modal self-supervision. Besides CMA, TCL introduces an intra-modal contrastive objective to provide complementary benefits in representation learning. To take advantage of localized and structural information from image and text input, TCL further maximizes the average MI between local regions of image/text and their global summary. To the best of our knowledge, ours is the first work that takes into account local structure information for multi-modality representation learning. Experimental evaluations show that our approach is competitive and achieves the new state of the art on various common downstream vision-language tasks such as image-text retrieval and visual question answering.

Similar Papers
  • Conference Article
  • Citations9

MAMO: Fine-Grained Vision-Language Representations Learning with Masked Multimodal Modeling

  • Jul 18, 2023
  • Zijia Zhao +5
  • Conference Article
  • Citations93

MixGen: A New Multi-Modal Data Augmentation

  • Jan 01, 2023
  • Xiaoshuai Hao +6
  • Book Chapter
  • Citations5

Switch-BERT: Learning to Model Multimodal Interactions by Switching Attention and Input

  • Jan 01, 2022
  • Qingpei Guo +2
  • Research Article
  • Citations460

Multimodal Intelligence: Representation Learning, Information Fusion, and Applications

  • Mar 01, 2020
  • IEEE Journal of Selected Topics in Signal Processing
  • Chao Zhang +3
  • Research Article
  • Citations1

Enhancing vision–language contrastive representation learning using domain knowledge

  • Sep 01, 2025
  • Computer Vision and Image Understanding
  • Xiaoyang Wei +2
  • Conference Article
  • Citations24

Exploiting Mutual Information for Substructure-aware Graph Representation Learning

  • Jul 01, 2020
  • Pengyang Wang +5
  • Research Article
  • Citations10

Robust visual question answering via semantic cross modal augmentation

  • Oct 16, 2023
  • Computer Vision and Image Understanding
  • Akib Mashrur +3
  • Research Article
  • Citations56

BridgeTower: Building Bridges between Encoders in Vision-Language Representation Learning

  • Jun 26, 2023
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Xiao Xu +5
  • Conference Article
  • Citations289

Contrastive Learning based Hybrid Networks for Long-Tailed Image Classification

  • Jun 01, 2021
  • Peng Wang +4
  • PDF
  • Research Article
  • Citations3

Clustering swap prediction for image-text pre-training

  • May 24, 2024
  • Scientific Reports
  • Sun Fayou +3
  • PDF
  • Research Article

Learning Multimodal Representations by Symmetrically Transferring Local Structures

  • Sep 13, 2020
  • Symmetry
  • Bin Dong +2
  • Conference Article
  • Citations102

Vector-Quantized Autoregressive Predictive Coding

  • Oct 25, 2020
  • Yu-An Chung +2
  • Research Article
  • Citations142

Vision-Language Pre-Training: Basics, Recent Advances, and Future Trends

  • Jan 01, 2022
  • Foundations and Trends® in Computer Graphics and Vision
  • Zhe Gan +5
  • Research Article
  • Citations2

A Novel Pretrained General-purpose Vision Language Model for the Vietnamese Language

  • May 10, 2024
  • ACM Transactions on Asian and Low-Resource Language Information Processing
  • Dinh Anh Vu +2
  • Research Article
  • Citations49

Structure-aware protein self-supervised learning

  • Apr 03, 2023
  • Bioinformatics
  • Can (Sam) Chen +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.