• Home
  • Search
  • An Entropy Clustering Approach for Assessing Visual Question Difficulty
  • Cite Icon2
  • https://doi.org/10.1109/access.2020.3022063Copy DOI Icon

An Entropy Clustering Approach for Assessing Visual Question Difficulty

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

We propose a novel approach to identify the difficulty of visual questions for Visual Question Answering (VQA) without direct supervision or annotations to the difficulty. Prior works have considered the diversity of ground-truth answers of human annotators. In contrast, we analyze the difficulty of visual questions based on the behavior of multiple different VQA models. We propose to cluster the entropy values of the predicted answer distributions obtained by three different models: a baseline method that takes as input images and questions, and two variants that take as input images only and questions only. We use a simple k-means to cluster the visual questions of the VQA v2 validation set. Then we use state-of-the-art methods to determine the accuracy and the entropy of the answer distributions for each cluster. A benefit of the proposed method is that no annotation of the difficulty is required, because the accuracy of each cluster reflects the difficulty of visual questions that belong to it. Our approach can identify clusters of difficult visual questions that are not answered correctly by state-of-the-art methods. Detailed analysis on the VQA v2 dataset reveals that 1) all methods show poor performances on the most difficult cluster (about 10% accuracy), 2) as the cluster difficulty increases, the answers predicted by the different methods begin to differ, and 3) the values of cluster entropy are highly correlated with the cluster accuracy. We show that our approach has the advantage of being able to assess the difficulty of visual questions without ground-truth (i.e., the test set of VQA v2) by assigning them to one of the clusters. We expect that this can stimulate the development of novel directions of research and new algorithms.

Loading PDF

Similar Papers
  • Research Article

Visual Question Answering Model Using Transformer-Based Multi-Head Cross Attention Network with Optimized Gold Rush Octave Convolution Network

  • Oct 28, 2025
  • International Journal of Image and Graphics
  • J Jinu Sophia +2
  • Conference Article
  • Citations90

Multi-Modality Latent Interaction Network for Visual Question Answering

  • Oct 01, 2019
  • Gao Peng +4
  • PDF
  • Conference Article
  • Citations37

The Promise of Premise: Harnessing Question Premises in Visual Question Answering

  • Jan 01, 2017
  • Aroma Mahendru +4
  • PDF
  • Conference Article
  • Citations33

WeaQA: Weak Supervision via Captions for Visual Question Answering

  • Jan 01, 2021
  • Pratyay Banerjee +3
  • Research Article
  • Citations7

Attacking VQA Systems via Adversarial Background Noise

  • Aug 01, 2020
  • IEEE Transactions on Emerging Topics in Computational Intelligence
  • Akshay Chaturvedi +1
  • Conference Article
  • Citations2

MVQAS

  • Oct 26, 2021
  • Haoyue Bai +3
  • Research Article
  • Citations10

Robust visual question answering via semantic cross modal augmentation

  • Oct 16, 2023
  • Computer Vision and Image Understanding
  • Akib Mashrur +3
  • Research Article
  • Citations16

Test-Time Model Adaptation for Visual Question Answering With Debiased Self-Supervisions

  • Jan 01, 2024
  • IEEE Transactions on Multimedia
  • Zhiquan Wen +5
  • Research Article
  • Citations41

Multitask Learning for Visual Question Answering.

  • Mar 01, 2023
  • IEEE Transactions on Neural Networks and Learning Systems
  • Jie Ma +5
  • Conference Article
  • Citations115

Overcoming Language Priors with Self-supervised Learning for Visual Question Answering

  • Jul 01, 2020
  • Xi Zhu +5
  • Research Article
  • Citations2

VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Chun-Mei Feng +7
  • Research Article
  • Citations34

Safety compliance checking of construction behaviors using visual question answering

  • Oct 07, 2022
  • Automation in Construction
  • Yuexiong Ding +2
  • Research Article
  • Citations6

Counting-based visual question answering with serial cascaded attention deep learning

  • Aug 02, 2023
  • Pattern Recognition
  • Tesfayee Meshuwelde +1
  • Conference Article
  • Citations2

Estimating Viewed Images with Natural Language Question Answering from fMRI Data

  • Mar 01, 2020
  • Saya Takada +3
  • Conference Article
  • Citations20

From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question Answering

  • Oct 17, 2021
  • Mingrui Lao +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.