Post

New research · Otolaryngology (ENT)
Frontiers in digital health · 1w
AI / informaticsFrontiers in digital health · 2026

A deep learning approach for acoustic-based identification of muscle tension dysphonia and spasmodic dysphonia.

Zhou Zhou, Yuan Cheng, Qingyi Ren … Pingjiang Ge
Read paper
Otolaryngology (ENT)AI / informatics

Deep learning distinguished healthy from disordered voices with 89.5% accuracy on audio alone

A deep learning approach for acoustic-based identification of muscle tension dysphonia and spasmodic dysphonia.

Zhou Zhou … Pingjiang Ge
Frontiers in digital health · 2026
Background

PROBLEM: Differentiating between spasmodic dysphonia (SD), a neurological disorder, and muscle tension dysphonia (muscle tension dysphonia), a behavioral voice disorder, based on auditory perception alone is a common clinical challenge.

Purpose

This study aims to develop and validate an artificial intelligence (artificial intelligence) model based on deep learning to automatically differentiate between healthy voices, SD, and muscle tension dysphonia using only voice audio recordings, and to compare its diagnostic performance against human experts.

Methods

A retrospective analysis was conducted on 1,597 voice samples (595 healthy, 471 muscle tension dysphonia, 531 SD).

n = 1,597 voice samples
89.5%
Results
89.5%
Model correctly sorted healthy versus disordered voices in 89.5% of recordings
n = 1,597 voice samples
More results

For the more complex ternary classification, the model attained an accuracy of 71.6%, with class-specific AUCs of 0.957 (Healthy), 0.731 (muscle tension dysphonia), and 0.855 (SD).

This performance surpassed that of human experts, who achieved average accuracies of 78.2% in binary classification and 60.6% in ternary classification on the same test set.

“
Conclusion · 1 of 2

The deep learning model trained on a Mandarin pathological voice dataset achieves favorable performance in distinguishing SD from muscle tension dysphonia using only voice audio signals, with classification results comparable to those of experienced clinical specialists.

Conclusion · 2 of 2

This technology serves as a promising objective auxiliary tool for the preliminary screening and differential diagnosis of voice disorders.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
≥ 95%
share of readers who felt risks, benefits, and other treatment options were well explained
AI / Informatics
44.7%
ear, nose, and throat doctors correctly spotted the writer less than half the time
AI / Informatics
4.50
expert doctors rated human-written reviews' overall scientific quality highest, above both AI chatbots
Cohort Study
5%
patients whose frontal sinus infection came back after surgery, so most stayed infection-free
Cohort Study
100%
all confirmed vocal cord paralysis cases needed surgery to restore voice, none healed on their own
Guideline
“Lingual frenectomy may improve maternal pain during breastfeeding and may be an option in selected cases of phonetic alterations.”
Study
65.4% reduction
fewer patients needed surgery to stop bleeding after tonsil removal
Cross-sectional
21.1%
fear of illness was the most common workplace challenge for ear, nose, and throat staff