Comparative Analysis of Time-Frequency Representations for Pediatric Respiratory Sound Classification Using Deep Learning
Respiratory sound classification has emerged as a promising non-invasive and scalable tool for the early detection of respiratory disorders. While most previous studies have relied on a single feature extraction method-such as Melfrequency cepstral coefficients (MFCC), Mel-spectrograms, or Short-Time Fourier Transform (STFT)-this study provides a comprehensive comparative analysis of these three approaches, evaluated individually, in pairwise combinations, and in a combined three-method configuration. Experiments are conducted on the SPRSound dataset, a pediatric respiratory sound database comprising 6,656 annotated events from seven unbalanced classes. Three architectures are assessed under identical preprocessing and augmentation strategies: a custom convolutional neural network (CNN), and two pre-trained models (VGG16 and InceptionV3) fine-tuned via transfer learning. Results show that STFT consistently delivers the highest performance for CNN and VGG16 models, while MFCC achieves the best accuracy with InceptionV3. Specifically, VGG16 with STFT attained 93.43% accuracy (score: 0.9545), whereas InceptionV3 with MFCC achieved the top performance within its architecture. These findings highlight the importance of aligning feature extraction techniques with model architecture and provide a systematic benchmark for SPRSound-based respiratory sound classification.
Read more