- Conference Article
- 10.1109/iisec69317.2026.11418501
Cross-Lingual Detection of Synthetic Speech Using Acoustic and Deep Features
- Feb 05, 2026
- Furkan Karaman + 2 more +2
Advances in synthetic speech generation pose significant risks to voice communication security. This study proposes an empirical framework for synthetic speech detection with an emphasis on cross-language robustness. Acoustic and prosodic features, including MFCC, CQCC, pitch, jitter, and shimmer, are evaluated using classical machine learning methods and deep learning models. Experiments are conducted on the FoR (monolingual) and ODSS (multilingual) benchmark datasets using stratified 80%–20% train–test splits and stratified 5-fold cross-validation. On FoR, a CNN achieves 99.2% accuracy with a 0.75% Equal Error Rate (EER). On ODSS, ResNet reaches 96.5% accuracy and 3.72% EER under the percentage split. Additionally, Random Forest attains 100% accuracy and 0% EER in 5-fold cross-validation on ODSS, indicating strong feature separability while suggesting potential overfitting effects. The results demonstrate that the proposed hybrid features and deep architectures provide high generalizability across different languages and synthesis techniques.
Read more