• Home
  • Search
  • Pre-trained transformer-based language models for Sundanese
  • Cite Icon18
  • https://doi.org/10.1186/s40537-022-00590-7Copy DOI Icon

Pre-trained transformer-based language models for Sundanese

Show More
  • Abstract
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The Sundanese language has over 32 million speakers worldwide, but the language has reaped little to no benefits from the recent advances in natural language understanding. Like other low-resource languages, the only alternative is to fine-tune existing multilingual models. In this paper, we pre-trained three monolingual Transformer-based language models on Sundanese data. When evaluated on a downstream text classification task, we found that most of our monolingual models outperformed larger multilingual models despite the smaller overall pre-training data. In the subsequent analyses, our models benefited strongly from the Sundanese pre-training corpus size and do not exhibit socially biased behavior. We released our models for other researchers and practitioners to use.

Loading PDF

Similar Papers
  • Research Article
  • Citations14

Fine-Tuned Understanding: Enhancing Social Bot Detection With Transformer-Based Classification

  • Jan 01, 2024
  • IEEE Access
  • Amine Sallah +6
  • Research Article
  • Citations77

Bangla-BERT: Transformer-Based Efficient Model for Transfer Learning and Language Understanding

  • Jan 01, 2022
  • IEEE Access
  • M Kowsher +5
  • Research Article
  • Citations27

Enhancing Transformer-based language models with commonsense representations for knowledge-driven machine comprehension

  • Mar 06, 2021
  • Knowledge-Based Systems
  • Ronghan Li +5
  • Research Article
  • Citations4

Cross-lingual dependency parsing for a language with a unique script

  • Sep 09, 2024
  • Natural Language Processing
  • He Zhou +2
  • PDF
  • Conference Article
  • Citations2

Probing Structured Pruning on Multilingual Pre-trained Models: Settings, Algorithms, and Efficiency

  • Jan 01, 2022
  • Yanyang Li +5
  • Video Transcripts

Robust Transfer Learning with Pretrained Language Models through Adapters

  • Aug 01, 2021
  • Underline Science Inc.
  • Wenjuan Han +2
  • PDF
  • Research Article
  • Citations10

AGI-P: A Gender Identification Framework for Authorship Analysis Using Customized Fine-Tuning of Multilingual Language Model

  • Jan 01, 2024
  • IEEE Access
  • Raheem Sarwar +6
  • Research Article
  • Citations7

The Generalization and Robustness of Transformer-Based Language Models on Commonsense Reasoning

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Ke Shen
  • PDF
  • Research Article
  • Citations26

Predicting Generalized Anxiety Disorder From Impromptu Speech Transcripts Using Context-Aware Transformer-Based Neural Networks: Model Evaluation Study

  • Mar 28, 2023
  • JMIR Mental Health
  • Bazen Gashaw Teferra +1
  • Conference Article
  • Citations9

On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code

  • Nov 30, 2023
  • Martin Weyssow +4
  • Conference Article

Pre-trained Language Models as Prior Knowledge for Playing Text-based Games

  • May 09, 2022
  • Ishika Singh +2
  • Conference Article

RBG-AI: Benefits of Multilingual Language Models for Low-Resource Languages

  • Jan 01, 2025
  • Barathi Ganesh Hb +1
  • PDF
  • Conference Article
  • Citations16

Time-Stamped Language Model: Teaching Language Models to Understand The Flow of Events

  • Jan 01, 2021
  • Hossein Rajaby Faghihi +1
  • Research Article

PromptSED: An evolving topic-enhanced prompting framework for incremental social event detection.

  • Nov 01, 2025
  • Neural networks : the official journal of the International Neural Network Society
  • Xiaoyan Yu +8
  • PDF
  • Research Article
  • Citations33

Depression Classification From Tweets Using Small Deep Transfer Learning Language Models

  • Jan 01, 2022
  • IEEE Access
  • Muhammad Rizwan +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.