Post

New research · Otolaryngology (ENT)
Frontiers in digital health · 1w
AI / informaticsFrontiers in digital health · 2026

A deep learning approach for acoustic-based identification of muscle tension dysphonia and spasmodic dysphonia.

Zhou Zhou, Yuan Cheng, Qingyi Ren … Pingjiang Ge
Read paper
Otolaryngology (ENT)AI / informatics

Deep learning distinguished healthy from disordered voices with 89.5% accuracy on audio alone

A deep learning approach for acoustic-based identification of muscle tension dysphonia and spasmodic dysphonia.

Zhou Zhou … Pingjiang Ge
Frontiers in digital health · 2026
Background

PROBLEM: Differentiating between spasmodic dysphonia (SD), a neurological disorder, and muscle tension dysphonia (muscle tension dysphonia), a behavioral voice disorder, based on auditory perception alone is a common clinical challenge.

Purpose

This study aims to develop and validate an artificial intelligence (artificial intelligence) model based on deep learning to automatically differentiate between healthy voices, SD, and muscle tension dysphonia using only voice audio recordings, and to compare its diagnostic performance against human experts.

Methods

A retrospective analysis was conducted on 1,597 voice samples (595 healthy, 471 muscle tension dysphonia, 531 SD).

n = 1,597 voice samples
89.5%
Results
89.5%
Model correctly sorted healthy versus disordered voices in 89.5% of recordings
n = 1,597 voice samples
More results

For the more complex ternary classification, the model attained an accuracy of 71.6%, with class-specific AUCs of 0.957 (Healthy), 0.731 (muscle tension dysphonia), and 0.855 (SD).

This performance surpassed that of human experts, who achieved average accuracies of 78.2% in binary classification and 60.6% in ternary classification on the same test set.

“
Conclusion · 1 of 2

The deep learning model trained on a Mandarin pathological voice dataset achieves favorable performance in distinguishing SD from muscle tension dysphonia using only voice audio signals, with classification results comparable to those of experienced clinical specialists.

Conclusion · 2 of 2

This technology serves as a promising objective auxiliary tool for the preliminary screening and differential diagnosis of voice disorders.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Randomized Trial
20.0
points lower score on a nasal symptom questionnaire with surgery than sprays
Animal / Preclinical
0%
no deep learning studies were tested with patients in real healthcare settings
Randomized Trial
-1·60
dupilumab improved nasal polyp severity more than omalizumab
Cross-sectional
21.1%
fear of illness was the most common workplace challenge for ear, nose, and throat staff
Guideline
“Lingual frenectomy may improve maternal pain during breastfeeding and may be an option in selected cases of phonetic alterations.”
AI / Informatics
4.50
expert doctors rated human-written reviews' overall scientific quality highest, above both AI chatbots
Cohort Study
51.2%
over half of children with a breathing tube experienced problems later on
Observational
β = 0.075
worse emotional well-being was associated with higher self-stigma