• Home
  • Search
  • A Masked Reference Token Supervision-Based Iterative Visual-Language Framework for Robust Visual Grounding
  • Cite Icon3
  • https://doi.org/10.1109/tcsvt.2024.3452418Copy DOI Icon

A Masked Reference Token Supervision-Based Iterative Visual-Language Framework for Robust Visual Grounding

  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Visual Grounding (VG) has become a prominent task in recent years, achieving significant advancements with the development of detection and vision transformers. However, existing VG methods struggle to handle the effects of inaccurate or irrelevant textual descriptions, tending to generate false-alarm objects. Moreover, existing methods fail to capture fine-grained features, accurate localization, and comprehensive context understanding from the whole image and textual descriptions. To address these issues, we propose an Iterative Robust Visual Grounding (IR-VG) framework with Multi-stage False-alarm Sensitive Decoder (MFSD) to prevent the generation of false-alarm objects when presented with inaccurate expressions. The framework introduces Masked Reference based Centerpoint Supervision (MRCS) and Iterative Multi-level Vision-language Fusion (IMVF) for enhancing the accuracy of localization and better visual-language alignment. To investigate the elements that affect VG robustness further, we release a robust VG benchmark with 24,000 instances and we also provide a detailed classification of false-alarm according to different parts of speech. Extensive experiments on existing state-of-the-art (SOTA) VG methods and foundation models have proven that it is difficult to handle the robustness of VG by existing models. Even foundation models, which have been pre-trained with a large amount of data, have difficulty to understand inaccurate language descriptions. Our IR-VG can handle false-alarm issues in robust VG well and achieve new SOTA results on the newly proposed robust VG datasets. Ablation studies and visualization experiments demonstrate the effectiveness of the proposed components. Moreover, the proposed framework is also verified effective on five regular VG datasets. Codes and models will be publicly at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/cv516Buaa/IR-VG</uri>.

Similar Papers
  • Research Article
  • Citations10

Language-guided Residual Graph Attention Network and Data Augmentation for Visual Grounding

  • Aug 24, 2023
  • ACM Transactions on Multimedia Computing, Communications, and Applications
  • Jia Wang +3
  • Research Article
  • Citations14

Visual Grounding With Joint Multimodal Representation and Interaction

  • Jan 01, 2023
  • IEEE Transactions on Instrumentation and Measurement
  • Hong Zhu +5
  • Conference Article
  • Citations79

3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point Clouds

  • Jun 01, 2022
  • Daigang Cai +4
  • Research Article

End-to-end Visual Grounding Based on Query Text Guidance and Multi-stage Reasoning

  • Feb 01, 2024
  • 電腦學刊
  • Chao Wang Chao Wang +5
  • Research Article
  • Citations19

When multiple instance learning meets foundation models: Advancing histological whole slide image analysis.

  • Apr 01, 2025
  • Medical image analysis
  • Hongming Xu +9
  • Research Article

Boosting foundation models for rare eye disease diagnosis via a multimodal text-to-image generative framework.

  • Mar 24, 2026
  • NPJ digital medicine
  • Ruoyu Chen +8
  • Conference Article

Towards A New Era of Geo-Foundation Models: Expert-Guided Multimodal Alignment and Geospatial Context Awareness

  • Nov 03, 2025
  • Ting Han +4
  • Book Chapter
  • Citations63

UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling

  • Jan 01, 2022
  • Zhengyuan Yang +7
  • Research Article

BioFuse: an embedding fusion framework for biomedical foundation models.

  • Mar 18, 2026
  • PloS one
  • Mirza Nasir Hossain +1
  • Research Article

GeoBiked: a dataset with geometric features and automated labeling techniques to advance deep generative models in engineering design

  • Jun 11, 2025
  • Engineering Computations
  • Phillip Mueller +2
  • Research Article
  • Citations142

Vision-Language Pre-Training: Basics, Recent Advances, and Future Trends

  • Jan 01, 2022
  • Foundations and Trends® in Computer Graphics and Vision
  • Zhe Gan +5
  • PDF
  • Research Article
  • Citations410

DECIPHER: harnessing local sequence context to improve protein multiple sequence alignment

  • Oct 06, 2015
  • BMC Bioinformatics
  • Erik S Wright
  • Research Article

Is Federated Learning Still Alive in the Foundation Model Era?

  • May 20, 2024
  • Proceedings of the AAAI Symposium Series
  • Nathalie Baracaldo
  • Conference Article
  • Citations2

Language Driven Image Editing via Transformers

  • Oct 01, 2022
  • Rodrigo Santos +2
  • Research Article
  • Citations3

The Impact of AI Foundation models on the future of digital engineering for logistics and supply Chain

  • Dec 01, 2025
  • Digital Engineering
  • Bernardo Nicoletti +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.