• Home
  • Search
  • Examining Sentiment Analysis for Low-Resource Languages with Data Augmentation Techniques
  • Cite Icon2
  • https://doi.org/10.3390/eng5040152Copy DOI Icon

Examining Sentiment Analysis for Low-Resource Languages with Data Augmentation Techniques

  • Nov 7, 2024
  • Eng
  • Gaurish Thakkar +2 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

This investigation investigates the influence of a variety of data augmentation techniques on sentiment analysis in low-resource languages, with a particular emphasis on Bulgarian, Croatian, Slovak, and Slovene. The following primary research topic is addressed: is it possible to improve sentiment analysis efficacy in low-resource languages through data augmentation? Our sub-questions look at how different augmentation methods affect performance, how effective WordNet-based augmentation is compared to other methods, and whether lemma-based augmentation techniques can be used, especially for Croatian sentiment tasks. The sentiment-labelled evaluations in the selected languages are included in our data sources, which were curated with additional annotations to standardise labels and mitigate ambiguities. Our findings show that techniques like replacing words with synonyms, masked language model (MLM)-based generation, and permuting and combining sentences can only make training datasets slightly bigger. However, they provide limited improvements in model accuracy for low-resource language sentiment classification. WordNet-based techniques, in particular, exhibit a marginally superior performance compared to other methods; however, they fail to substantially improve classification scores. From a practical perspective, this study emphasises that conventional augmentation techniques may require refinement to address the complex linguistic features that are inherent to low-resource languages, particularly in mixed-sentiment and context-rich instances. Theoretically, our results indicate that future research should concentrate on the development of augmentation strategies that introduce novel syntactic structures rather than solely relying on lexical variations, as current models may not effectively leverage synonymic or lemmatised data. These insights emphasise the nuanced requirements for meaningful data augmentation in low-resource linguistic settings and contribute to the advancement of sentiment analysis approaches.

Similar Papers
  • Research Article
  • Citations25

Improving the prediction of extreme wind speed events with generative data augmentation techniques

  • Nov 30, 2023
  • Renewable Energy
  • M Vega-Bayo +3
  • PDF
  • Research Article
  • Citations72

Evaluation of Data Augmentation Techniques for Facial Expression Recognition Systems

  • Nov 11, 2020
  • Electronics
  • Simone Porcu +2
  • Research Article

Self-Supervised Deep Learning Models for Low-Resource NLP Applications

  • Feb 26, 2026
  • International Journal of Research and Review in Applied Science, Humanities, and Technology
  • Pardeep Kaur
  • Conference Article

Benchmarking BoW, TF-IDF and Bert-Based Sentiment Analysis in Low-Resource Minang Language

  • Oct 09, 2025
  • Iip Permana +5
  • Research Article

Inter-machine harmonization of multicenter echocardiographic images for improvement of left ventricular ejection fraction prediction model

  • Oct 27, 2025
  • Scientific Reports
  • Ren Iwasaki +7
  • Conference Article
  • Citations23

PatchAugment: Local Neighborhood Augmentation in Point Cloud Classification

  • Oct 01, 2021
  • Shivanand Venkanna Sheshappanavar +2
  • Research Article

Deep Learning-Based Eye-Writing Recognition with Improved Preprocessing and Data Augmentation Techniques

  • Oct 13, 2025
  • Sensors (Basel, Switzerland)
  • Kota Suzuki +2
  • Conference Article
  • Citations37

Breast Cancer Detection Using GAN for Limited Labeled Dataset

  • Sep 25, 2020
  • Shrinivas D Desai +4
  • Research Article
  • Citations15

YOLOv8-Based Drone Detection: Performance Analysis and Optimization

  • Sep 17, 2024
  • Computers
  • Betul Yilmaz +1
  • Research Article
  • Citations81

Deep Adversarial Data Augmentation for Extremely Low Data Regimes

  • Jan 23, 2020
  • IEEE Transactions on Circuits and Systems for Video Technology
  • Xiaofeng Zhang +4
  • Research Article
  • Citations13

A Roman Urdu Corpus for sentiment analysis

  • Jun 18, 2024
  • The Computer Journal
  • Marwa Khan +3
  • Research Article
  • Citations1

Mathematical Problem Solving in Arabic: Assessing Large Language Models

  • Jan 01, 2024
  • Procedia Computer Science
  • Abeer Mahgoub +2
  • PDF
  • Research Article
  • Citations14

LiDA: Language-Independent Data Augmentation for Text Classification

  • Jan 01, 2023
  • IEEE Access
  • Yudianto Sujana +1
  • Conference Article
  • Citations11

Data Augmentation Model for Audio Signal Extraction

  • Aug 17, 2022
  • M Muthumari +3
  • Book Chapter
  • Citations10

Geometric Transformations-Based Medical Image Augmentation

  • Jan 01, 2023
  • S Kalaivani +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.