• Home
  • Search
  • Segmentation-Driven Deep Learning for Explainable and Automated Vocal Fold Disorder Classification
  • https://doi.org/10.1109/icecte69292.2026.11429171Copy DOI Icon

Segmentation-Driven Deep Learning for Explainable and Automated Vocal Fold Disorder Classification

  • Jan 29, 2026
  • Fahima Sharmin Hossain +2 more
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

The abnormalities of the vocal folds cause voice disorders that affect speech and communication. The traditional diagnostic tools such as laryngoscopy and high-speed video endoscopy are subjective and time consuming and may be affected by inter-observer variation. In order to overcome these difficulties, we suggest a deep-learning-based paradigm of automatic vocal-fold segmentation and diagnosis of disorder on the Emory Vocal Folds dataset (approximately 2.9k of annotated images). We use preprocessing (cleaning, normalization and augmentation, class balancing), semantic segmentation as U-Net, and feature extractions as state-of-the-art CNNs (ResNet, DenseNet, VGG16, MobileNet and EfficientNet). The extracted features were used with both the end-to-end CNN classification and traditional machine learning classifiers (SVM, Random Forest, Gradient Boosting). It was experimentally demonstrated that EfficientNet was the most accurate at validation (98%), as compared to ResNet, DenseNet, and VGG16. Classical models also showed good results with an accuracy of 97% with random forest and 96 with SVM. The interpretability of the models was guaranteed through Grad-CAM, Grad-CAM++, Integrated Gradients, and Occlusion Sensitivity, and it was confirmed that the predictions paid attention to the parts of the vocal folds that are clinically relevant. These results indicate that deep learning can offer a robust, interpretable, and non-invasive algorithm for the early detection and diagnosis of vocal fold disorders thereby providing a possibility of real-time clinical decision support and mobile health diagnosis.

Similar Papers
  • Research Article

Glottis Analysis Tools (Kist et al., 2021)

  • May 17, 2021
  • Figshare
  • Andreas M Kist +11
  • Research Article
  • Citations35

Efficacy of Videostroboscopy and High-Speed Videoendoscopy to Obtain Functional Outcomes From Perioperative Ratings in Patients With Vocal Fold Mass Lesions

  • Apr 17, 2019
  • Journal of Voice
  • Maria E Powell +6
  • Research Article
  • Citations5

Android malware detection using GIST based machine learning and deep learning techniques

  • Aug 01, 2024
  • Indonesian Journal of Electrical Engineering and Computer Science
  • Ponnuswamy Udayakumar +5
  • Research Article
  • Citations23

Spatial Segmentation for Laryngeal High-Speed Videoendoscopy in Connected Speech

  • Nov 27, 2020
  • Journal of Voice
  • Ahmed M Yousef +5
  • Research Article

Studying the glottal vibration onset and offset using laryngeal high-speed videoendoscopy in connected speech

  • Mar 01, 2023
  • The Journal of the Acoustical Society of America
  • Maryam Naghibolhosseini +3
  • Research Article
  • Citations54

Utility of Laryngeal High-speed Videoendoscopy in Clinical Voice Assessment

  • Jun 07, 2017
  • Journal of Voice
  • Stephanie R.C Zacharias +2
  • Research Article
  • Citations7

Laryngeal High-Speed Videoendoscopy with Laser Illumination: A Preliminary Report.

  • Sep 07, 2021
  • Otolaryngologia Polska
  • Jakub Malinowski +8
  • Research Article
  • Citations2

Voice Recovery in a Patient with Inhaled Laryngeal Burns.

  • Jan 01, 2019
  • Iranian journal of otorhinolaryngology
  • Geun-Hyo Kim +3
  • Research Article

Unilateral Vocal Fold Paralysis: Does the Paralyzed Vocal Fold Oscillate Slower or Faster Than the Healthy One?

  • May 01, 2025
  • Journal of voice : official journal of the Voice Foundation
  • Dominika Valášková +4
  • PDF
  • Research Article
  • Citations10

High-speed kymography identifies the immediate effects of voiced vibration in healthy vocal folds

  • Jan 01, 2013
  • International Archives of Otorhinolaryngology
  • María Dájer +5
  • Research Article
  • Citations3

OPTIMIZATION OF LUNG CANCER CLASSIFICATION METHOD USING EDA-BASED MACHINE LEARNING

  • Feb 28, 2023
  • Jurnal Sistem Informasi dan Ilmu Komputer Prima(JUSIKOM PRIMA)
  • Windania Purba +4
  • Research Article

Voice Research and Treatment in the Czech Republic

  • May 01, 2006
  • The ASHA Leader
  • Jan G Svec +1
  • Research Article

SHAP-Enhanced Machine Learning for Explainable Stroke Risk Prediction in Hypertensive Patients

  • Dec 31, 2025
  • Journal of Material Sciences and Engineering Technology
  • David Olanrewaju Akinwale +2
  • Research Article
  • Citations11

Wideostrobokimografia w praktyce foniatrycznej

  • May 01, 2012
  • Otolaryngologia Polska
  • Agata Szkiełkowska +4
  • Research Article
  • Citations1

Comparison of Machine Learning and Deep Learning Models Performance in predicting wind energy

  • Jul 21, 2025
  • EAI Endorsed Transactions on Energy Web
  • Saswati Rakshit +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.