Post

New research · Otolaryngology (ENT)
World journal of otorhinolaryngology - head and neck surgery · 1w
AI / informaticsWorld journal of otorhinolaryngology - head and neck surgery · 2026

Blinded by the Bot: Benchmarking GPT and Gemini Against Human Authors in Otolaryngology Reviews.

Sholem Hack, Rebecca Attal, Letizia Nitro … Masayoshi Takashima
Read paper
Otolaryngology (ENT)AI / informatics

Human-written otolaryngology reviews rated higher quality than ChatGPT and Gemini reviews

Blinded by the Bot: Benchmarking GPT and Gemini Against Human Authors in Otolaryngology Reviews.

Sholem Hack … Masayoshi Takashima
World journal of otorhinolaryngology - head and neck surgery · 2026
Purpose

To compare the quality of scientific review articles generated by two artificial intelligence systems, ChatGPT and Gemini, with those written by human authors in the field of otolaryngology.

Methods

Two otolaryngology topics, chronic rhinosinusitis and infantile subglottic hemangioma, were selected.

n = 10 manuscripts
Results

expert doctors rated human-written reviews' overall scientific quality highest, above both AI chatbots

Human-authored
4.55-poin
GPT-4.0
2.715-poin
Gemini 2.0
2.145-poin
More results

GPT-4.0 demonstrated moderate performance (2.71 ± 1.46 overall), while Gemini 2.0 scored lowest (2.14 ± 1.01).

Mixed-effects modeling demonstrated significant group differences across all domains ( p ≤ 0.008).

More results

Citation quality showed one of the largest between-group differences and strong reliability [intraclass correlation coefficients(2,1) = 0.68; intraclass correlation coefficients(2,7) = 0.94].

“
Conclusion · 1 of 3

GPT generated fluent, stylistically strong reviews but remained significantly inferior to human-authored manuscripts in analytical depth and citation integrity.

Conclusion · 2 of 3

Gemini 2.0 underperformed across all domains and demonstrated a substantial rate of fabricated citations.

Conclusion · 3 of 3

As large language models become integrated into academic workflows, transparent disclosure, structured fact-checking, and human oversight remain essential to safeguard scientific reliability.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Randomized Trial
20.0
points lower score on a nasal symptom questionnaire with surgery than sprays
Randomized Trial
-1·60
dupilumab improved nasal polyp severity more than omalizumab
Animal / Preclinical
0%
no deep learning studies were tested with patients in real healthcare settings
Cross-sectional
21.1%
fear of illness was the most common workplace challenge for ear, nose, and throat staff
Guideline
“Lingual frenectomy may improve maternal pain during breastfeeding and may be an option in selected cases of phonetic alterations.”
AI / Informatics
89.5%
Model correctly sorted healthy versus disordered voices in 89.5% of recordings
Cohort Study
51.2%
over half of children with a breathing tube experienced problems later on
Observational
β = 0.075
worse emotional well-being was associated with higher self-stigma