• Home
  • Search
  • Robust visual question answering via semantic cross modal augmentation
  • Cite Icon10
  • https://doi.org/10.1016/j.cviu.2023.103862Copy DOI Icon

Robust visual question answering via semantic cross modal augmentation

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Recent advances in vision-language models have resulted in improved accuracy in visual question answering (VQA) tasks. However, their robustness remains limited when faced with out-of-distribution data containing unanswerable questions. In this study, we first construct a simple randomised VQA dataset, incorporating unanswerable questions from the VQA v2 dataset, to evaluate the robustness of a state-of-the-art VQA model. Our findings reveal that the model struggles to predict the “unknown” answer or provides inaccurate responses with high confidence scores for irrelevant questions. To address this issue without retraining the large backbone models, we propose Cross Modal Augmentation (CMA), a model-agnostic, test-time-only, multi-modal semantic augmentation technique. CMA generates multiple semantically-consistent but heterogeneous instances from the visual and textual inputs, which are then fed to the model, and the predictions are combined to achieve a more robust output. We demonstrate that implementing CMA enables the VQA model to provide more reliable answers in scenarios involving unanswerable questions, and show that the approach is generalisable across different categories of pre-trained vision language models.

Similar Papers
  • Research Article

Visual Question Answering Model Using Transformer-Based Multi-Head Cross Attention Network with Optimized Gold Rush Octave Convolution Network

  • Oct 28, 2025
  • International Journal of Image and Graphics
  • J Jinu Sophia +2
  • PDF
  • Conference Article
  • Citations33

WeaQA: Weak Supervision via Captions for Visual Question Answering

  • Jan 01, 2021
  • Pratyay Banerjee +3
  • Book Chapter
  • Citations2

Convolutional Neural Networks-Based VQA Model

  • Jun 28, 2022
  • Himanshu Sharma +1
  • Research Article
  • Citations16

Test-Time Model Adaptation for Visual Question Answering With Debiased Self-Supervisions

  • Jan 01, 2024
  • IEEE Transactions on Multimedia
  • Zhiquan Wen +5
  • Conference Article
  • Citations2

MVQAS

  • Oct 26, 2021
  • Haoyue Bai +3
  • PDF
  • Conference Article
  • Citations37

The Promise of Premise: Harnessing Question Premises in Visual Question Answering

  • Jan 01, 2017
  • Aroma Mahendru +4
  • Conference Article
  • Citations115

Overcoming Language Priors with Self-supervised Learning for Visual Question Answering

  • Jul 01, 2020
  • Xi Zhu +5
  • Research Article
  • Citations2

VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Chun-Mei Feng +7
  • Research Article
  • Citations34

Safety compliance checking of construction behaviors using visual question answering

  • Oct 07, 2022
  • Automation in Construction
  • Yuexiong Ding +2
  • Research Article
  • Citations6

Counting-based visual question answering with serial cascaded attention deep learning

  • Aug 02, 2023
  • Pattern Recognition
  • Tesfayee Meshuwelde +1
  • Conference Article
  • Citations20

From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question Answering

  • Oct 17, 2021
  • Mingrui Lao +5
  • PDF
  • Conference Article
  • Citations3

EaSe: A Diagnostic Tool for VQA based on Answer Diversity

  • Jan 01, 2021
  • Shailza Jolly +2
  • PDF
  • Research Article
  • Citations14

An effective spatial relational reasoning networks for visual question answering.

  • Nov 28, 2022
  • PLOS ONE
  • Xiang Shen +4
  • Research Article
  • Citations64

Image captioning for effective use of language models in knowledge-based visual question answering

  • Aug 28, 2022
  • Expert Systems with Applications
  • Ander Salaberria +4
  • Research Article
  • Citations1

Towards Robust Visual Question Answering via Prompt-Driven Geometric Harmonization

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Yishu Liu +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.