• Home
  • Search
  • Thai Sentence Completeness Classification using Fine-Tuned WangchanBERTa
  • https://doi.org/10.52783/jisem.v10i34s.5821Copy DOI Icon

Thai Sentence Completeness Classification using Fine-Tuned WangchanBERTa

Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Sentence completeness classification plays a crucial role in various natural language processing (NLP) applications, including grammar checking, text auto-completion, and language assessment. This task becomes particularly challenging in Thai due to the language’s unique characteristics such as flexible word order, implicit subject omission, and the absence of explicit word boundaries. These linguistic properties make traditional rule-based and statistical approaches prone to errors when applied to Thai. To address these challenges, this research applies modern deep learning techniques, specifically leveraging pre-trained transformer models fine-tuned for Thai sentence completeness classification. This study introduces the use of WangchanBERTa, a Thai-specific adaptation of RoBERTa, pre-trained. Two thousand Thai sentences are created. There are one thousand complete sentences and one thousand incomplete sentences. Each sentence was manually labeled to ensure high data quality. Experimental results show that WangchanBERTa achieves an average accuracy of 99.65%, significantly outperforming mBERT, a popular multilingual baseline, which achieved only 95.82%. Notably, with a Tesla T4 GPU, WangchanBERTa required just 1 hour and 15 minutes to train across all folds, compared to mBERT’s 2 hours and 59 minutes. Additionally, WangchanBERTa’s performance was compared with XLM-R, a state-of-the-art multilingual model, which achieved a slightly higher accuracy of 99.90% but at the cost of higher computational requirements. The results emphasize the advantage of language-specific pretraining in capturing the linguistic nuances of Thai. This research highlights the importance of tailored transformer models for low-resource languages. By demonstrating that WangchanBERTa achieves near state-of- the-art performance with lower computational cost, this work provides a strong foundation for future Thai NLP research.

Similar Papers
  • PDF
  • Research Article

Review of Language Structures and NLP Techniques for Chinese, Japanese, and English

  • Nov 22, 2024
  • Applied and Computational Engineering
  • Jingxuan Du
  • Research Article
  • Citations1

Fine-Tuning QurSim on Monolingual and Multilingual Models for Semantic Search

  • Jan 23, 2025
  • Information
  • Tania Afzal +3
  • Research Article

Commentaries on Computer-based Language Assessment

  • Dec 22, 2012
  • Sae Rhim Oh +1
  • Research Article
  • Citations77

Bangla-BERT: Transformer-Based Efficient Model for Transfer Learning and Language Understanding

  • Jan 01, 2022
  • IEEE Access
  • M Kowsher +5
  • Research Article
  • Citations73

The promise of NLP and speech processing technologies in language assessment

  • Jul 01, 2010
  • Language Testing
  • Carol A Chapelle +1
  • PDF
  • Research Article
  • Citations11

Extracting Pulmonary Nodules and Nodule Characteristics from Radiology Reports of Lung Cancer Screening Patients Using Transformer Models

  • May 17, 2024
  • Journal of Healthcare Informatics Research
  • Shuang Yang +10
  • Research Article

Constructing and Evaluating a Japanese Entity Linking Model for Wikidata based on Pointer Network

  • Nov 01, 2024
  • Transactions of the Japanese Society for Artificial Intelligence
  • Yuki Sawamura +4
  • Conference Article
  • Citations3

Question Classification from Thai Sentences by Considering Word Context to Question Generation

  • Aug 04, 2022
  • Saranlita Chotirat +2
  • PDF
  • Research Article

Does learning from language family help? A case study on a low-resource question-answering task

  • Jun 03, 2024
  • Natural Language Processing
  • Hariom A Pandya +1
  • Research Article

Advancements in NLP for social media analytics

  • May 11, 2025
  • International Journal of Scientific World
  • Harem Hadi +1
  • Research Article
  • Citations1

HARNESSING THE USE OF ARTIFICIAL INTELLIGENCE IN LANGUAGE ASSESSMENT: A SYSTEMATIC COMPREHENSIVE REVIEW

  • Sep 30, 2023
  • TELL-US JOURNAL
  • Hidayatul Ummah Al-Imamah Al-Abbas +2
  • Research Article

Natural language processing for geriatric syndromes: a systematic review of methods, applications, and challenges.

  • Mar 12, 2026
  • BMC medical informatics and decision making
  • Fahrurrozi Rahman +8
  • Research Article

Processing Low-Resource Languages: A Review Of Challenges And Strategies For Inclusive NLP And Sustainable Environment

  • Sep 19, 2025
  • International Journal of Environmental Sciences
  • Diganta Baishya +3
  • Research Article
  • Citations1

Improving Part-of-Speech Tagging with Relative Positional Encoding in Transformer Models and Basic Rules

  • Mar 31, 2025
  • Indonesian Journal of Data and Science
  • Jerome Aondongu Achir +2
  • Book Chapter
  • Citations1

Automatic Question and Answer Generation from Thai Sentences

  • Jan 01, 2022
  • Saranlita Chotirat +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.