Post

New research · Otolaryngology (ENT)
World journal of otorhinolaryngology - head and neck surgery · 1w
AI / informaticsWorld journal of otorhinolaryngology - head and neck surgery · 2026

Blinded by the Bot: Benchmarking GPT and Gemini Against Human Authors in Otolaryngology Reviews.

Sholem Hack, Rebecca Attal, Letizia Nitro … Masayoshi Takashima
Read paper
Otolaryngology (ENT)AI / informatics

Human-written otolaryngology reviews rated higher quality than ChatGPT and Gemini reviews

Blinded by the Bot: Benchmarking GPT and Gemini Against Human Authors in Otolaryngology Reviews.

Sholem Hack … Masayoshi Takashima
World journal of otorhinolaryngology - head and neck surgery · 2026
Purpose

To compare the quality of scientific review articles generated by two artificial intelligence systems, ChatGPT and Gemini, with those written by human authors in the field of otolaryngology.

Methods

Two otolaryngology topics, chronic rhinosinusitis and infantile subglottic hemangioma, were selected.

n = 10 manuscripts
Results

expert doctors rated human-written reviews' overall scientific quality highest, above both AI chatbots

Human-authored
4.55-poin
GPT-4.0
2.715-poin
Gemini 2.0
2.145-poin
More results

GPT-4.0 demonstrated moderate performance (2.71 ± 1.46 overall), while Gemini 2.0 scored lowest (2.14 ± 1.01).

Mixed-effects modeling demonstrated significant group differences across all domains ( p ≤ 0.008).

More results

Citation quality showed one of the largest between-group differences and strong reliability [intraclass correlation coefficients(2,1) = 0.68; intraclass correlation coefficients(2,7) = 0.94].

“
Conclusion · 1 of 3

GPT generated fluent, stylistically strong reviews but remained significantly inferior to human-authored manuscripts in analytical depth and citation integrity.

Conclusion · 2 of 3

Gemini 2.0 underperformed across all domains and demonstrated a substantial rate of fabricated citations.

Conclusion · 3 of 3

As large language models become integrated into academic workflows, transparent disclosure, structured fact-checking, and human oversight remain essential to safeguard scientific reliability.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
89.5%
Model correctly sorted healthy versus disordered voices in 89.5% of recordings
AI / Informatics
≥ 95%
share of readers who felt risks, benefits, and other treatment options were well explained
AI / Informatics
44.7%
ear, nose, and throat doctors correctly spotted the writer less than half the time
Cohort Study
5%
patients whose frontal sinus infection came back after surgery, so most stayed infection-free
Cohort Study
100%
all confirmed vocal cord paralysis cases needed surgery to restore voice, none healed on their own
Guideline
“Lingual frenectomy may improve maternal pain during breastfeeding and may be an option in selected cases of phonetic alterations.”
Study
65.4% reduction
fewer patients needed surgery to stop bleeding after tonsil removal
Cross-sectional
21.1%
fear of illness was the most common workplace challenge for ear, nose, and throat staff