- Conference Article
1
- 10.1145/3697467.3697659
Multi-level Semantic Feature Fusion Text Classification Model Based on BERT
- Aug 09, 2024
- Chen Feng + 2 more +2
With the rapid development of information technology and the Internet, text information on various media platforms, such as blogs, microblogs, and news, has exploded. In order to process these text data more efficiently, text classification technology has gradually become a research focus. In view of the common problems of incomplete structure and sparse features in short texts, this paper designs a short text classification model with multi-level semantic feature fusion to effectively extract key word features and sentence features from short texts. The model first uses BERT for preliminary semantic feature extraction, and then combines the first-layer Bi-GRU and the word-level self-attention mechanism to enhance the weight of important semantic information in the text. Next, the second-layer Bi-GRU and the sentence-level self-attention mechanism are used to achieve feature fusion at the word vector and sentence vector levels. The maximum pooling strategy of CNN is used to screen the best features and predict the category probability. In addition, the loss function is adjusted, a threshold is introduced to screen and continuously train samples that are difficult to classify, and an adaptive adjustment factor is added in the back propagation to adaptively attenuate the learning rate and accelerate the model fitting process. In the experiment of 20 Newsgroups dataset, compared with the baseline model, the classification accuracy of this model increased by 3.51% to 88.74%, which effectively proved the effectiveness of this model in improving the classification accuracy of Chinese short texts.
Read more