• Home
  • Search
  • Tibetan Sentence Boundaries Automatic Disambiguation Based on Bidirectional Encoder Representations from Transformers on Byte Pair Encoding Word Cutting Method
  • Cite Icon3
  • https://doi.org/10.3390/app14072989Copy DOI Icon

Tibetan Sentence Boundaries Automatic Disambiguation Based on Bidirectional Encoder Representations from Transformers on Byte Pair Encoding Word Cutting Method

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Sentence Boundary Disambiguation (SBD) is crucial for building datasets for tasks such as machine translation, syntactic analysis, and semantic analysis. Currently, most automatic sentence segmentation in Tibetan adopts the methods of rule-based and statistical learning, as well as the combination of the two, which have high requirements on the corpus and the linguistic foundation of the researchers and are more costly to annotate manually. In this study, we explore Tibetan SBD using deep learning technology. Initially, we analyze Tibetan characteristics and various subword techniques, selecting Byte Pair Encoding (BPE) and Sentencepiece (SP) for text segmentation and training the Bidirectional Encoder Representations from Transformers (BERT) pre-trained language model. Secondly, we studied the Tibetan SBD based on different BERT pre-trained language models, which mainly learns the ambiguity of the shad (“།”) in different positions in modern Tibetan texts and determines through the model whether the shad (“།”) in the texts has the function of segmenting sentences. Meanwhile, this study introduces four models, BERT-CNN, BERT-RNN, BERT-RCNN, and BERT-DPCNN, based on the BERT model for performance comparison. Finally, to verify the performance of the pre-trained language models on the SBD task, this study conducts SBD experiments on both the publicly available Tibetan pre-trained language model TiBERT and the multilingual pre-trained language model (Multi-BERT). The experimental results show that the F1 score of the BERT (BPE) model trained in this study reaches 95.32% on 465,669 Tibetan sentences, nearly five percentage points higher than BERT (SP) and Multi-BERT. The SBD method based on pre-trained language models in this study lays the foundation for establishing datasets for the later tasks of Tibetan pre-training, summary extraction, and machine translation.

Loading PDF

Similar Papers
  • Research Article
  • Citations43

Augmenting commit classification by using fine-grained source code changes and a pre-trained deep neural language model

  • Mar 10, 2021
  • Information and Software Technology
  • Lobna Ghadhab +3
  • PDF
  • Research Article
  • Citations25

Investigating the Impact of Prompt Engineering on the Performance of Large Language Models for Standardizing Obstetric Diagnosis Text: Comparative Study

  • Feb 08, 2024
  • JMIR Formative Research
  • Lei Wang +7
  • Research Article

Knowledge enhancement BERT based on domain dictionary mask

  • Apr 21, 2023
  • Journal of High Speed Networks
  • Xianglin Cao +2
  • Research Article
  • Citations3

(Retracted) News image text classification algorithm with bidirectional encoder representations from transformers model

  • Sep 13, 2022
  • Journal of Electronic Imaging
  • Zhan Shi +3
  • Research Article

RESEARCH OF THE PROCESS OF VISUAL ART TRANSMISSION IN MUSIC AND THE CREATION OF COLLECTIONS FOR PEOPLE WITH VISUAL IMPAIRMENTS

  • Apr 02, 2025
  • Municipal economy of cities
  • I Karavan +4
  • PDF
  • Research Article
  • Citations6

Identification and Impact Analysis of Family History of Psychiatric Disorder in Mood Disorder Patients With Pretrained Language Model

  • May 20, 2022
  • Frontiers in Psychiatry
  • Cheng Wan +7
  • PDF
  • Research Article

Comparison of active learning algorithms in classifying head computed tomography reports using bidirectional encoder representations from transformers

  • Jan 08, 2025
  • International Journal of Computer Assisted Radiology and Surgery
  • Tomohiro Wataya +12
  • Research Article
  • Citations1

Application of the Bidirectional Encoder Representations from Transformers Model for Predicting the Abbreviated Injury Scale in Patients with Trauma: Algorithm Development and Validation Study

  • May 29, 2025
  • JMIR Formative Research
  • Jun Tang +5
  • PDF
  • Research Article
  • Citations141

BERT Models for Arabic Text Classification: A Systematic Review

  • Jun 04, 2022
  • Applied Sciences
  • Ali Saleh Alammary
  • Research Article

Leveraging Contextual Information in Biomedical Named Entity Recognition: Determining the Optimal Depth of Context

  • Nov 18, 2025
  • International Journal on Artificial Intelligence Tools
  • S M Archana +1
  • Research Article
  • Citations13

Deep Neural Networks in Natural Language Processing for Classifying Requirements by Origin and Functionality: An Application of BERT in System Requirements

  • Nov 13, 2023
  • Journal of Mechanical Design
  • Jesse Mullis +3
  • Dissertation

Using large language models for analyzing drug label data

  • Mar 01, 2024
  • Taha Valizadehaslani +2
  • PDF
  • Research Article
  • Citations26

BERT-Kgly: A Bidirectional Encoder Representations From Transformers (BERT)-Based Model for Predicting Lysine Glycation Site for Homo sapiens.

  • Feb 18, 2022
  • Frontiers in Bioinformatics
  • Yinbo Liu +5
  • Research Article
  • Citations14

Fine-Tuned Understanding: Enhancing Social Bot Detection With Transformer-Based Classification

  • Jan 01, 2024
  • IEEE Access
  • Amine Sallah +6
  • Research Article

Comparative Analysis of Indonesian Pre-trained BERT Models for the Extractive Question Answering Task on an Indonesian-Translated SQuAD Dataset

  • Mar 11, 2026
  • MATRIK : Jurnal Manajemen, Teknik Informatika dan Rekayasa Komputer
  • Fattah Al Ilmi Suhendra +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.