• Home
  • Search
  • Prediction and feature selection of low birth weight using machine learning algorithms
  • Cite Icon14
  • https://doi.org/10.1186/s41043-024-00647-8Copy DOI Icon

Prediction and feature selection of low birth weight using machine learning algorithms

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Background and aimsThe birth weight of a newborn is a crucial factor that affects their overall health and future well-being. Low birth weight (LBW) is a widespread global issue, which the World Health Organization defines as weighing less than 2,500 g. LBW can have severe negative consequences on an individual’s health, including neonatal mortality and various health concerns throughout their life. To address this problem, this study has been conducted using BDHS 2017–2018 data to uncover important aspects of LBW using a variety of machine learning (ML) approaches and to determine the best feature selection technique and best predictive ML model.MethodsTo pick out the key features, the Boruta algorithm and wrapper method were used. Logistic Regression (LR) used as traditional method and several machine learning classifiers were then used, including, DT (Decision Tree), SVM (Support Vector Machine), NB (Naïve Bayes), RF (Random Forest), XGBoost (eXtreme Gradient Boosting), and AdaBoost (Adaptive Boosting), to determine the best model for predicting LBW. The model’s performance was evaluated based on the specificity, sensitivity, accuracy, F1 score and AUC value.ResultsResult shows, Boruta algorithm identifies eleven significant features including respondent’s age, highest education level, educational attainment, wealth index, age at first birth, weight, height, BMI, age at first sexual intercourse, birth order number, and whether the child is a twin. Incorporating Boruta algorithm’s significant features, the performance of traditional LR and ML methods including DT, SVM, NB, RF, XGBoost, and AB were evaluated where LR, had a specificity, sensitivity, accuracy and F1 score of 0.85, 0.5, 85.15% and 0.915. While the ML methods DT, SVM, NB, RF, XGBoost, and AB model’s respective accuracy values were 85.35%, 85.15%, 84.54%, 81.18%, and 84.41%. Based on the specificity, sensitivity, accuracy, F1 score and AUC, RF (specificity = 0.99, sensitivity = 0.58, accuracy = 85.86%, F1 score = 0.9243, AUC = 0.549) outperformed the other methods. Both the classical (LR) and machine learning (ML) models’ performance has improved dramatically when important characteristics are extracted using the wrapper method. The LR method identified five significant features with a specificity, sensitivity, accuracy and F1 score of 0.87, 0.33, 87.12% and 0.9309. The region, whether the infant is a twin, and cesarean delivery were the three key features discovered by the DT and RF models, which were implemented using the wrapper technique. All three models had the identical F1 score of 0.9318. However, “child is twin” was recognized as a significant feature by the SVM, NB, and AB models, with an F1 score of 0.9315. Ultimately, with an F1 score of 0.9315, the XGBoost model recognized “child is twin” and “age at first sex” as relevant features. Random Forest again beat the other approaches in this instance.ConclusionsThe study reveals Wrapper method as the optimal feature selection technique. The ML method outperforms traditional methods, with Random Forest (RF) being the most effective predictive model for Low-Birth-Weight prediction. The study suggests that policymakers in Bangladesh can mitigate low birth weight newborns by considering identified risk factors.

Similar Papers
  • Research Article
  • Citations5

Application of an interpretable machine learning method to predict the risk of death during hospitalization in patients with acute myocardial infarction combined with diabetes mellitus

  • Apr 07, 2025
  • Acta Cardiologica
  • Zhijun Bu +12
  • Research Article
  • Citations5

Comparative study of different machine learning models in landslide susceptibility assessment: A case study of Conghua District, Guangzhou, China

  • Jan 01, 2024
  • China Geology
  • Ao Zhang +10
  • Research Article
  • Citations1

Breast Cancer Detection Analysis Using Different Machine Learning Techniques: South Iraq Case Study

  • Feb 26, 2025
  • International Journal of Prognostics and Health Management
  • Salma Abdulbaki Mahmood +5
  • Research Article

Construction and preliminary validation of machine learning predictive models for cervical cancer screening based on human DNA methylation

  • Feb 23, 2025
  • Zhonghua zhong liu za zhi [Chinese journal of oncology]
  • Y Yang +9
  • Research Article
  • Citations17

Sex determination through maxillary dental arch and skeletal base measurements using machine learning

  • Aug 30, 2024
  • Head & Face Medicine
  • Cristiano Miranda De Araujo +9
  • Research Article
  • Citations27

Parameter importance assessment improves efficacy of machine learning methods for predicting snow avalanche sites in Leh-Manali Highway, India

  • Jun 29, 2021
  • Science of the Total Environment
  • Anuj Tiwari +2
  • PDF
  • Research Article
  • Citations5

Prediction of Lumbar Drainage-Related Meningitis Based on Supervised Machine Learning Algorithms

  • Jun 28, 2022
  • Frontiers in Public Health
  • Peng Wang +6
  • PDF
  • Research Article
  • Citations32

The Machine Learning Model for Distinguishing Pathological Subtypes of Non-Small Cell Lung Cancer.

  • May 26, 2022
  • Frontiers in Oncology
  • Hongyue Zhao +9
  • Book Chapter

From Optimal Model to Web Application

  • May 09, 2025
  • Van Dai Pham +3
  • PDF
  • Research Article
  • Citations5

Establishment of a risk prediction model for olfactory disorders in patients with transnasal pituitary tumors by machine learning

  • May 31, 2024
  • Scientific Reports
  • Min Chen +8
  • Research Article
  • Citations3

Failure in Stock Price Prediction: A Comparison bettwen the Curve-Shape-Feature and Non-Curve-Shape-Feature Modes of Existing Machine Learning Algorithms

  • Dec 03, 2021
  • INTERNATIONAL JOURNAL OF COMPUTERS COMMUNICATIONS & CONTROL
  • Ping Zhang +5
  • Conference Article
  • Citations9

Statistical evaluation for quality of experience prediction based on quality of service parameters

  • May 01, 2016
  • Sana Aroussi +1
  • Research Article
  • Citations52

Applying data mining techniques to classify patients with suspected hepatitis C virus infection

  • Jan 10, 2022
  • Intelligent Medicine
  • Reza Safdari +3
  • Research Article
  • Citations13

A Dataset Centric Feature Selection and Stacked Model to Detect Breast Cancer

  • Aug 08, 2021
  • International Journal of Intelligent Systems and Applications
  • Avijit Kumar Chaudhuri +2
  • Research Article
  • Citations1

Predicting the risk of threatened abortion using machine learning methods: a comparative study

  • Aug 30, 2025
  • BMC Pregnancy and Childbirth
  • Zhenning Zhu +8
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.