A deep learning approach for acoustic-based identification of muscle tension dysphonia and spasmodic dysphonia.
Deep learning distinguished healthy from disordered voices with 89.5% accuracy on audio alone
A deep learning approach for acoustic-based identification of muscle tension dysphonia and spasmodic dysphonia.
PROBLEM: Differentiating between spasmodic dysphonia (SD), a neurological disorder, and muscle tension dysphonia (muscle tension dysphonia), a behavioral voice disorder, based on auditory perception alone is a common clinical challenge.
This study aims to develop and validate an artificial intelligence (artificial intelligence) model based on deep learning to automatically differentiate between healthy voices, SD, and muscle tension dysphonia using only voice audio recordings, and to compare its diagnostic performance against human experts.
A retrospective analysis was conducted on 1,597 voice samples (595 healthy, 471 muscle tension dysphonia, 531 SD).
For the more complex ternary classification, the model attained an accuracy of 71.6%, with class-specific AUCs of 0.957 (Healthy), 0.731 (muscle tension dysphonia), and 0.855 (SD).
This performance surpassed that of human experts, who achieved average accuracies of 78.2% in binary classification and 60.6% in ternary classification on the same test set.
The deep learning model trained on a Mandarin pathological voice dataset achieves favorable performance in distinguishing SD from muscle tension dysphonia using only voice audio signals, with classification results comparable to those of experienced clinical specialists.
This technology serves as a promising objective auxiliary tool for the preliminary screening and differential diagnosis of voice disorders.