Post

New research · Ophthalmology
International ophthalmology · 3d
AI / informaticsInternational ophthalmology · 2026

Comparative performance of chatgpt and gemini in diagnostic classification and clinical reasoning for open-angle glaucoma: a standardized scenario-based study.

Zhewen Zhang, Siyu Lu, Zhenqiang Xu … Yan Liang
Read paper
OphthalmologyAI / informatics

ChatGPT (GPT-5.3) scored higher than Gemini on clinical reasoning for glaucoma cases

Comparative performance of chatgpt and gemini in diagnostic classification and clinical reasoning for open-angle glaucoma: a standardized scenario-based study.

Zhewen Zhang … Yan Liang
International ophthalmology · 2026
Purpose

To evaluate differences in performance between two large language models (large language models), GPT-5.3 and Gemini 2.5 Pro, in diagnostic classification and clinical reasoning for primary open-angle glaucoma (primary open-angle glaucoma).

Methods

We compared the diagnostic accuracy and classification consistency (Cohen's κ) of the two models.

Results
4.4
vs. 3.9
on reasoning quality checks, GPT-5.3 scored better than Gemini
More results

Overall diagnostic accuracy was 85.4% for GPT-5.3 and 75.0% for Gemini (P = 0.306).

For classification consistency, κ values were 0.675 for GPT-5.3 and 0.628 for Gemini.

Error pattern analysis indicated that Gemini was more prone to overdiagnosis and reasoning inconsistency, whereas GPT-5.3 was relatively conservative.

More results

Both models had low rates of unsafe outputs, though Gemini showed a slightly higher proportion.

“
Conclusion · 1 of 2

ChatGPT and Gemini both demonstrate certain capabilities in diagnosing primary open-angle glaucoma, but their stability on borderline cases remains limited. Comparatively, GPT-5.3 shows higher consistency and more stable reasoning patterns.

Conclusion · 2 of 2

The application of large language models in ophthalmic diagnostic support still requires cautious evaluation.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
0·962 for OCT
MerMED-FM was very accurate at diagnosing diseases using eye scans
AI / Informatics
0.98
Artificial intelligence's eyelid-height measurements matched doctors' manual measurements almost perfectly
AI / Informatics
92.1%
combining eye scans and photos correctly told benign from cancerous lesions apart nearly every time
Cohort Study
28.6%
recurred locally in more than 1 in 4 patients over years of follow-up
Cohort Study
-0.395
excision group's post-op eyelid fullness score was lower - greater improvement
Cohort Study
86.7%
of infants probed after 12 months still had unresolved tear duct blockage
Observational
16%
eyelid tissue in rosacea patients showed less of this key repair-signaling protein inside cell nuclei
Cohort Study
25%
about 1 in 4 treated cases had symptoms return after improving