• Home
  • Search
  • Not a statistical accident: framing bias as a semiotic property of image datasets
  • https://doi.org/10.1007/s00146-026-02948-4Copy DOI Icon

Not a statistical accident: framing bias as a semiotic property of image datasets

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Abstract Despite being a well-studied problem, bias continues to affect modern Computer Vision (CV) datasets, models and systems, including Generative AI models that have been found to replicate and amplify harmful biases and stereotypes. In this work, we focus on framing bias : a type of bias that arises from the way images convey meaning. We introduce a semiotic perspective that treats image datasets as texts and argues that framing bias is not a statistical accident due to sampling noise or imbalance, but an inherent property of image datasets as meaning-making systems. We show that co-occurrence of visual elements is the primary mechanism through which image datasets frame concepts, and that not only do individual images contribute to dataset framing through co-occurrence, but dataset-level framing also shapes the interpretation of individual images. As a consequence, every image dataset embodies a point of view and cannot be fully unbiased . We illustrate these claims through a case study of the Visual Genome dataset, revealing a framing of human activities that over-represents leisure, systematically excludes everyday labour, and is characterised by a predominantly American/Western viewpoint. Finally, we reflect on the epistemological implications of analysing image datasets, highlighting the mediated and interpretive nature of knowledge produced through AI systems, and propose revisionism and source criticism as epistemological paradigms to address these issues.

Similar Papers
  • Research Article
  • Citations51

The informed sampler: A discriminative approach to Bayesian inference in generative computer vision models

  • May 24, 2015
  • Computer Vision and Image Understanding
  • Varun Jampani +3
  • Research Article

SELF-SUPERVISED VISION TRANSFORMERS FOR CROSS-MODAL LEARNING (REVIEW)

  • Jan 01, 2025
  • Computer Design Systems. Theory and Practice
  • Olena Stankevych +1
  • Conference Article
  • Citations21

Enhancing Fairness in Face Detection in Computer Vision Systems by Demographic Bias Mitigation

  • Jul 26, 2022
  • Yu Yang +8
  • Conference Article

Advancements in Defense Mechanisms against Adversarial Attacks in Computer Vision

  • Dec 23, 2024
  • Satish S Banait +5
  • Front Matter
  • Citations2

Guest Editors' Introduction: Special Section on Higher Order Graphical Models in Computer Vision.

  • Jul 01, 2015
  • IEEE transactions on pattern analysis and machine intelligence
  • Karteek Alahari +4
  • Research Article
  • Citations46

DeepSolar++: Understanding residential solar adoption trajectories with computer vision and technology diffusion models

  • Nov 01, 2022
  • Joule
  • Zhecheng Wang +4
  • Single Book
  • Citations587

Handbook of Mathematical Models in Computer Vision

  • Jan 01, 2006
  • Nikos Paragios +2
  • Research Article
  • Citations1

412 A model generalization study in localizing indoor cows with COw LOcalization (COLO) dataset.

  • Oct 04, 2025
  • Journal of Animal Science
  • C P James Chen +2
  • Research Article
  • Citations31

Automated seed identification with computer vision: challenges and opportunities

  • Oct 17, 2022
  • Seed Science and Technology
  • Liang Zhao +2
  • Book Chapter
  • Citations77

Transparency and Trust in Human-AI-Interaction: The Role of Model-Agnostic Explanations in Computer Vision-Based Decision Support

  • Jan 01, 2020
  • Christian Meske +1
  • Research Article

HOW NVIDIA GPUS POWER THE WORLD MOST ADVANCED AI MODELS

  • May 05, 2024
  • International Journal of Computer Science and Engineering Research and Development
  • Srinivas Kola
  • Research Article

Comprehensive image dataset of flexible pavement: Alligator cracks and edge-breaks from national highway (N6) of urban areas

  • Jan 30, 2026
  • Data in Brief
  • Md Nayem Hossain +5
  • Peer Review Report

Decision letter: THINGS-data, a multimodal collection of large-scale datasets for investigating object representations in human brain and behavior

  • Oct 26, 2022
  • Talia Konkle +1
  • Conference Article
  • Citations33

Multiclass Colorectal Cancer Histology Images Classification Using Vision Transformers

  • Dec 05, 2021
  • Magdy Abd-Elghany Zeid +2
  • Conference Article
  • Citations1

<title>Qualitative approach to medical image databases</title>

  • Jul 01, 1991
  • Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
  • Yves J Bizais +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.