- Conference Article
- 10.1109/icecmsn68058.2025.11382966
Sentence Classification in Myanmar Language using FastText Embeddings and BERT-based Features
- Nov 24, 2025
- Kyault Kyault Khaing + 2 more +2
Sentence classification is presently utilized in various fields because of its dependability, scalability, and efficiency. The grammatical structure of the Myanmar language presents specific difficulties for sentence classification. The lack of spaces between words complicates the identification of words and syllables during feature extraction. This characteristic significantly affects model performance and preprocessing, unlike in languages such as English. Several challenges must be addressed when performing classification modeling with actual data. The primary objective of sentence classification is to identify the type of subjective information present in the source material. A dataset for sentence classification in the Myanmar language has been developed using data collected from websites and social media platforms. Currently, 112K sentences have been gathered and labeled with class identifiers. The classification models employed for categorizing the input sentences include Feedforward Neural Network (FNN) and Bidirectional Long Short-Term Memory Networks (BiLSTM). To improve the effectiveness of sentence classification, BERT-based features and fastText embeddings are utilized. Both feature types demonstrate strong individual performance and work effectively with FNN and BiLSTM. The effectiveness of these features is also evident when they are combined in both FNN and BiLSTM. This paper focuses on validating the dataset and evaluating the impact of these features on sentence classification in the Myanmar language. Using the combined features, the FNN model reaches an average F1-score of 97%, while the BiLSTM model achieves an average F1-score of 98%.
Read more