• Home
  • Search
  • Training Dataset and Dictionary Sizes Matter in BERT Models: The Case of Baltic Languages
  • Open Access IconOpen Access
  • Cite Icon6
  • https://doi.org/10.1007/978-3-031-16500-9_14Copy DOI Icon

Training Dataset and Dictionary Sizes Matter in BERT Models: The Case of Baltic Languages

  • Jan 1, 2022
  • Matej Ulčar +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Abstract Large pretrained masked language models have become state-of-the-art solutions for many NLP problems. While studies have shown that monolingual models produce better results than multilingual models, the training datasets must be sufficiently large. We trained a trilingual LitLat BERT-like model for Lithuanian, Latvian, and English, and a monolingual Est-RoBERTa model for Estonian. We evaluate their performance on four downstream tasks: named entity recognition, dependency parsing, part-of-speech tagging, and word analogy. To analyze the importance of focusing on a single language and the importance of a large training set, we compare created models with existing monolingual and multilingual BERT models for Estonian, Latvian, and Lithuanian. The results show that the newly created LitLat BERT and Est-RoBERTa models improve the results of existing models on all tested tasks in most situations.KeywordsNatural language processingBERTTransformersEstonianLatvianLithuanian

Similar Papers
  • Research Article
  • Citations77

Bangla-BERT: Transformer-Based Efficient Model for Transfer Learning and Language Understanding

  • Jan 01, 2022
  • IEEE Access
  • M Kowsher +5
  • Research Article
  • Citations1

Fine-Tuning QurSim on Monolingual and Multilingual Models for Semantic Search

  • Jan 23, 2025
  • Information
  • Tania Afzal +3
  • Video Transcripts

When is BERT Multilingual? Isolating Crucial Ingredients for Cross-lingual Transfer

  • Jun 27, 2022
  • Underline Science Inc.
  • Partha Talukdar +2
  • PDF
  • Conference Article
  • Citations2

FiSSA at SemEval-2020 Task 9: Fine-tuned for Feelings

  • Jan 01, 2020
  • Bertelt Braaksma +4
  • Book Chapter
  • Citations9

A Study on the Impact of Intradomain Finetuning of Deep Language Models for Legal Named Entity Recognition in Portuguese

  • Jan 01, 2020
  • Luiz Henrique Bonifacio +3
  • Research Article
  • Citations4

Cross-lingual dependency parsing for a language with a unique script

  • Sep 09, 2024
  • Natural Language Processing
  • He Zhou +2
  • PDF
  • Research Article
  • Citations18

Pre-trained transformer-based language models for Sundanese

  • Apr 13, 2022
  • Journal of Big Data
  • Wilson Wongso +2
  • Research Article

Cross-lingual Training for Multiple-Choice Question Answering

  • Oct 22, 2020
  • Procesamiento Del Lenguaje Natural
  • Guillermo Echegoyen +2
  • Conference Article
  • Citations21

Czert – Czech BERT-like Model for Language Representation

  • Jan 01, 2021
  • Jakub Sido +5
  • PDF
  • Conference Article
  • Citations5

Grapheme-to-Phoneme Conversion with a Multilingual Transformer Model

  • Jan 01, 2020
  • Omnia Elsaadany +1
  • PDF
  • Conference Article
  • Citations3

Evaluating Cross-Lingual Transfer Learning Approaches in Multilingual Conversational Agent Models

  • Jan 01, 2020
  • Lizhen Tan +1
  • PDF
  • Conference Article
  • Citations2

Probing Structured Pruning on Multilingual Pre-trained Models: Settings, Algorithms, and Efficiency

  • Jan 01, 2022
  • Yanyang Li +5
  • PDF
  • Conference Article
  • Citations39

Universal Sentence Representation Learning with Conditional Masked Language Model

  • Jan 01, 2021
  • Ziyi Yang +4
  • Video Transcripts

MRAT-SQL+GAP: A Portuguese Text-to-SQL Transformer

  • Nov 16, 2021
  • Underline Science Inc.
  • Fábio Gagliardi Cozman +1
  • PDF
  • Conference Article
  • Citations243

Emerging Cross-lingual Structure in Pretrained Language Models

  • Jan 01, 2020
  • Alexis Conneau +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.