• Home
  • Search
  • Zero-shot Referring Image Segmentation Augmented by Multimodal Large Language Models
  • https://doi.org/10.1145/3788108.3788502Copy DOI Icon

Zero-shot Referring Image Segmentation Augmented by Multimodal Large Language Models

  • Nov 28, 2025
  • Yongqiu Huang +5 more
Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Zero shot learning brings new challenges to referring image segmentation, a typical application to find a segmentation mask given a referring expression. Most researches borrow generalization of multi-modal model, such as Contrastive Language Image Pre-training(CLIP), to realize zero shot inference. However, CLIP also suffers the limitation of insensitive to direction and local information, which is incapable of fine-grained region-text matching. To overcome this limitation, we propose a new framework, to leverage multimodal large language models (MLLM) into segmentation. Specifically, augmentations are considered from two aspects. One is to align local image information to comprehensive linguistic information, under the hypothesis that the textual description guiding local target object in image. Comprehensive linguistic information here are considered from local, global and multi-modal generated description. The other one is to enhance the robustness by blurring and captions from multi-modal model. Comprehensive experiments on three datasets prove the effectiveness of proposed approach.

Similar Papers
  • Research Article

Large-scale evaluation of multimodal large language models for pneumothorax detection.

  • Feb 12, 2026
  • Diagnostic and interventional radiology (Ankara, Turkey)
  • Hamza Eren Güzel +2
  • Research Article

Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation

  • Nov 19, 2025
  • ACM Transactions on Multimedia Computing, Communications, and Applications
  • Xiaobo Yang +1
  • Conference Article

Enhancing Sentiment Analysis with Multimodal Large Language Models

  • May 23, 2025
  • Thresa Jeniffer J +4
  • Research Article

TinnitusLLM: A Multimodal Large Language Model Framework for Tinnitus Diagnosis Through EEG-fMRI Fusion Learning.

  • Jan 01, 2026
  • IEEE journal of biomedical and health informatics
  • Yipeng Du +9
  • Research Article

Diagnostic accuracy of large language models in the classification of superior labial frenulum attachments.

  • Dec 10, 2025
  • Odontology
  • Mehmet Gümüş Kanmaz +1
  • Research Article
  • Citations22

VIGC: Visual Instruction Generation and Correction

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Bin Wang +10
  • Conference Article

Words Over Pixels? Rethinking Vision in Multimodal Large Language Models

  • Aug 01, 2024
  • Anubhooti Jain +2
  • Book Chapter

A Novel Level Set Model Based on Local Information

  • Jan 01, 2010
  • Hai Min +2
  • Research Article

The Power of Multimodality in Multimodal Large Language Models, Unimodal ChatGPT 5.0, and Human Clinical Experts on a Wound Care Certification Examination: Cross-Sectional Comparative Study.

  • Apr 27, 2026
  • JMIR formative research
  • Mete Ucdal +5
  • Research Article
  • Citations8

Medical MLLM Is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language Models

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Xijie Huang +8
  • Conference Article

Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models

  • Dec 08, 2025
  • Kimia Ehsani +1
  • Research Article
  • Citations15

Embodied Intelligence in Mining: Leveraging Multi-Modal Large Language Models for Autonomous Driving in Mines

  • May 01, 2024
  • IEEE Transactions on Intelligent Vehicles
  • Luxi Li +9
  • Research Article
  • Citations20

Mini-Gemini: Mining the Potential of Multi-Modality Vision Language Models.

  • Mar 01, 2026
  • IEEE transactions on pattern analysis and machine intelligence
  • Yanwei Li +7
  • Research Article
  • Citations15

Predicting Content Similarity via Multimodal Modeling for Video-In-Video Advertising

  • Mar 17, 2020
  • IEEE Transactions on Circuits and Systems for Video Technology
  • Xue Song +2
  • Research Article
  • Citations1

Orthodontic Biomechanical Reasoning with Multimodal Language Models: Performance and Clinical Utility

  • Oct 27, 2025
  • Bioengineering
  • Arda Arısan +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.