- Conference Article
- 10.1109/icoa66896.2025.11236878
Predicting industrial accident causes from text descriptions: a Machine Learning-Based approach
- Oct 16, 2025
- Kawtar Benderouach + 3 more +3
Automatic identification of accident causes from textual reports is a major challenge for improving risk management in industrial systems. This article proposes an approach based on automatic natural language processing (ANLP) and machine learning to predict accident causes from textual descriptions. A database of real accident reports was used, containing textual summaries (SummaryCase) and their associated causes (Cause). After text pre-processing including cleaning, tokenization and vectorization (TF-IDF and Bag of Words), several supervised learning models were tested, including random forests (RF), logistic regression (LR), support vector machines (SVM), Decision tree (DT), Naive Bayes (NB) as well as XGBoost (Extreme Gradient Boosting). The performance of the model was examined in terms of the accuracy, precision, recall and F1-score. As a result those work demonstrate that an XGBoost combining with SMOTE (Synthetic Minority Over-Sampling Technique) are more effective when compared to traditional models especially in the situation of long and complex text description. The results of this study have proved the effectiveness of the application of machine learning techniques to the proactive analysis of accident causes and are expected to contribute to the adaptation of these approaches into the realm of industrial risk management systems.
Read more