• Home
  • Search
  • Analyzing Bias in Large Language Models: A Quantitative Study Using Sentiment and Demographic Metrics
  • https://doi.org/10.58723/ijaaiml.v2i2.411Copy DOI Icon

Analyzing Bias in Large Language Models: A Quantitative Study Using Sentiment and Demographic Metrics

  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Background of study: The widespread adoption of Large Language Models (LLMs) raises concerns about biases that affect fairness and credibility. As LLMs affect areas such as recruitment and customer service, systematic quantitative analysis is essential to identify and mitigate these biases.Aims and scope of paper: This research investigates demographic bias in LLM quantitatively by analyzing sentiment polarity scores across different demographic categories. The goal is to provide a statistically confirmed analysis of sentiment bias and propose mitigation methods, focusing on GPT-4, LLaMA-2, Claude, and BLOOM.Methods: Quantitative analysis was performed on GPT-4, LLaMA-2, Claude, and BLOOM using sentiment and demographic data. Sentiment polarity assessments for gender and racial/ethnic groups were obtained with VADER and TextBlob. Demographic Disparity Score, ANOVA, and Cohen's Kappa assessed the significance and appropriateness of bias. Inter-rater reliability between automated tools and human annotators was also evaluated.Result: Sentiment bias was found in all models, varying by gender and race, particularly in GPT-4 and Claude. Sentiment scores were consistently higher for queries pertaining to females than those pertaining to males across all models, with GPT-4 and Claude showing the largest differences. Claude also showed racial sentiment alignment, favoring queries relating to white people over black people. ANOVA confirmed statistically significant sentiment variation by demographics across all models. High inter-rater reliability validated the sentiment analysis.Conclusion: This study shows demographic bias in GPT-4, LLaMA-2, Claude, and BLOOM, with different sentiment trends across demographic classifications. The models showed more positive sentiment for female questions and a trend towards certain racial groups. These findings indicate an embedded bias in the training data, which raises ethical concerns. Identifying and addressing these biases is critical to ensuring fairness and credibility in real-world LLM applications.

Similar Papers
  • Research Article

Evaluating gpt-4 for zero-shot classification of bleeding and clotting events: Can large language models serve as second reviewers?

  • Nov 03, 2025
  • Blood
  • Samantha Rizzo +5
  • Research Article

Abstract 4369198: Performance of Large Language Models in Analyzing Common Hypertension Scenarios in Clinical Practice

  • Nov 04, 2025
  • Circulation
  • Jaleh Zand +7
  • Research Article
  • Citations4

Extracting epilepsy-related information from unstructured clinic letters using large language models.

  • Jul 10, 2025
  • Epilepsia
  • Shichao Fang +7
  • Research Article

Artificial intelligence for immunotherapy response assessment in lung cancer using PET-CT reports.

  • Jun 01, 2025
  • Journal of Clinical Oncology
  • Ozden Altundag +10
  • Research Article
  • Citations30

Do ChatGPT and Gemini Provide Appropriate Recommendations for Pediatric Orthopaedic Conditions?

  • Aug 22, 2024
  • Journal of pediatric orthopedics
  • Sean Pirkle +2
  • Research Article

TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models

  • Oct 28, 2025
  • ACM Transactions on Software Engineering and Methodology
  • Florian Tambon +4
  • Research Article

Comparing guideline adherence and readability: Artificial intelligence with deep learning versus specialized physicians in peripheral artery disease management.

  • Dec 18, 2025
  • Vascular medicine (London, England)
  • Alfredo Verastegui +9
  • Research Article
  • Citations3

Evaluation and Bias Analysis of Large Language Models in Generating Synthetic Electronic Health Records: Comparative Study.

  • May 12, 2025
  • Journal of medical Internet research
  • Ruochen Huang +11
  • Research Article

Validating Radiology Artificial Intelligence Model Performance on Photon-Counting CT Images Using Large Language Models for Ground Truth Extraction.

  • Mar 01, 2026
  • Journal of the American College of Radiology : JACR
  • Yee Seng Ng +8
  • Research Article

Diagnostic Performance of ChatGPT-o1 and DeepSeek-V3 in Expert-Validated Simulated Ear Nose and Throat Scenarios: A Comparative Accuracy Study

  • Mar 26, 2026
  • European Journal of Rhinology and Allergy
  • Nazlım Hilal Taraf +7
  • Research Article

Large language models for toxicity extraction in oncology trials: A real-world benchmark in prostate radiotherapy.

  • Mar 01, 2026
  • Radiotherapy and oncology : journal of the European Society for Therapeutic Radiology and Oncology
  • Federico Mastroleo +9
  • Research Article

Comparison of artificial intelligence and multidisciplinary team recommendations in the management of colorectal cancer liver metastases.

  • Feb 04, 2026
  • Scientific reports
  • Mustafa Yılmaz +6
  • Research Article

Integrating Large Language Models with Robotics for Naturalistic Human–Robot Communication

  • Mar 31, 2025
  • International Journal of Advanced and Innovative Research (IJAIR)
  • Mohammed Ali
  • Research Article
  • Citations5

The Performance of ChatGPT-4 and Gemini Ultra 1.0 for Quality Assurance Review in Emergency Medical Services Chest Pain Calls

  • Jul 11, 2024
  • Prehospital Emergency Care
  • Graham Brant-Zawadzki +6
  • Research Article
  • Citations2

Large Language Models Using Clinical Text in Pediatrics

  • Mar 02, 2026
  • JAMA Network Open
  • Tracy Huang +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.