• Home
  • Search
  • FusionSAM: Visual Multi-Modal Learning with Segment Anything Model
  • Cite Icon2
  • https://doi.org/10.1145/3711896.3736973Copy DOI Icon

FusionSAM: Visual Multi-Modal Learning with Segment Anything Model

  • Aug 3, 2025
  • Daixun Li +7 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Multimodal image fusion and semantic segmentation are critical for autonomous driving. Despite advancements, current models often struggle with segmenting densely packed elements due to a lack of comprehensive fusion features for guidance during training. While the Segment Anything Model (SAM) allows precise control during fine-tuning through its flexible prompting encoder, its potential remains largely unexplored in the context of multimodal segmentation for natural images. In this paper, we introduce SAM into multimodal image segmentation for the first time, proposing a novel framework that combines Latent Space Token Generation (LSTG) and Fusion Mask Prompting (FMP) modules. This approach transforms the training methodology for multimodal segmentation from a traditional black-box approach to a controllable, prompt-based mechanism. Specifically, we obtain latent space features for both modalities through vector quantization and embed them into a cross-attention-based inter-domain fusion module to establish long-range dependencies between modalities. We then use these comprehensive fusion features as prompts to guide precise pixel-level segmentation. Extensive experiments on multiple public datasets demonstrate that our method significantly outperforms SAM and SAM2 in multimodal autonomous driving scenarios, achieving an average improvement of 4.1% over the state-of-the-art method in segmentation mIoU, and the performance is also optimized in other multi-modal visual scenes.

Similar Papers
  • Research Article
  • Citations20

A review of deep learning approaches for multimodal image segmentation of liver cancer.

  • Oct 07, 2024
  • Journal of applied clinical medical physics
  • Chaopeng Wu +10
  • Research Article
  • Citations2

Distilling Hierarchical Knowledge from Multimodal Fusion for Unimodal Image Segmentation

  • Jan 01, 2025
  • IEEE Transactions on Circuits and Systems for Video Technology
  • Yujia Sun +6
  • Research Article
  • Citations12

A modality-collaborative convolution and transformer hybrid network for unpaired multi-modal medical image segmentation with limited annotations.

  • Mar 15, 2023
  • Medical Physics
  • Hong Liu +6
  • Research Article
  • Citations15

Cross-modality synthesis aiding lung tumor segmentation on multi-modal MRI images

  • Mar 21, 2022
  • Biomedical Signal Processing and Control
  • Jiaxin Li +5
  • Book Chapter
  • Citations3

Chapter 4 - Robust watermarking algorithm based on multimodal medical image fusion

  • Jan 01, 2024
  • Data Fusion Techniques and Applications for Smart Healthcare
  • Om Prakash Singh +4
  • PDF
  • Research Article
  • Citations11

A Novel Multi-Modality Image Simultaneous Denoising and Fusion Method Based on Sparse Representation

  • Oct 13, 2021
  • Computers
  • Guanqiu Qi +4
  • Research Article
  • Citations159

An image quality enhancement scheme employing adolescent identity search algorithm in the NSST domain for multimodal medical image fusion

  • Feb 19, 2021
  • Biomedical Signal Processing and Control
  • Jais Jose +6
  • Research Article
  • Citations12

Pancreatic cancer segmentation in unregistered multi-parametric MRI with adversarial learning and multi-scale supervision

  • Oct 05, 2021
  • Neurocomputing
  • Jun Li +4
  • Research Article

Multimodal image fusion–guided microvascular decompression for hemifacial spasm: a comparative clinical study

  • Jun 01, 2026
  • Interdisciplinary Neurosurgery
  • Shuaida Man +3
  • Conference Article

A Dimension Hybrid Framework for Multimodal Medical Image Segmentation

  • Dec 10, 2021
  • Zichen Su +1
  • PDF
  • Research Article
  • Citations3

Development of a software system for surgical robots based on multimodal image fusion: study protocol.

  • Jun 06, 2024
  • Frontiers in surgery
  • Shuo Yuan +7
  • Book Chapter
  • Citations6

Performance Analysis of VGG19 Deep Learning Network Based Brain Image Fusion

  • Dec 01, 2020
  • Vijayarajan Rajangam +3
  • Conference Article

Multi-modal Image Fusion based Anatomical Shape Model for Low-contrast Anterior Visual Pathway and Medial Rectus Muscle Segmentation in CT Images

  • Oct 23, 2019
  • Guoyu Hu +5
  • Research Article
  • Citations14

MOFA: A novel dataset for Multi-modal Image Fusion Applications

  • Mar 23, 2023
  • Information Fusion
  • Kaihua Xiao +3
  • Research Article
  • Citations1

MMFNet: A Mamba-Based Multimodal Fusion Network for Remote Sensing Image Semantic Segmentation.

  • Oct 08, 2025
  • Sensors (Basel, Switzerland)
  • Jingting Qiu +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.