Post

New research · Ophthalmology
International ophthalmology · 3d
AI / informaticsInternational ophthalmology · 2026

Comparative performance of chatgpt and gemini in diagnostic classification and clinical reasoning for open-angle glaucoma: a standardized scenario-based study.

Zhewen Zhang, Siyu Lu, Zhenqiang Xu … Yan Liang
Read paper
OphthalmologyAI / informatics

ChatGPT (GPT-5.3) scored higher than Gemini on clinical reasoning for glaucoma cases

Comparative performance of chatgpt and gemini in diagnostic classification and clinical reasoning for open-angle glaucoma: a standardized scenario-based study.

Zhewen Zhang … Yan Liang
International ophthalmology · 2026
Purpose

To evaluate differences in performance between two large language models (large language models), GPT-5.3 and Gemini 2.5 Pro, in diagnostic classification and clinical reasoning for primary open-angle glaucoma (primary open-angle glaucoma).

Methods

We compared the diagnostic accuracy and classification consistency (Cohen's κ) of the two models.

Results
4.4
vs. 3.9
on reasoning quality checks, GPT-5.3 scored better than Gemini
More results

Overall diagnostic accuracy was 85.4% for GPT-5.3 and 75.0% for Gemini (P = 0.306).

For classification consistency, κ values were 0.675 for GPT-5.3 and 0.628 for Gemini.

Error pattern analysis indicated that Gemini was more prone to overdiagnosis and reasoning inconsistency, whereas GPT-5.3 was relatively conservative.

More results

Both models had low rates of unsafe outputs, though Gemini showed a slightly higher proportion.

“
Conclusion · 1 of 2

ChatGPT and Gemini both demonstrate certain capabilities in diagnosing primary open-angle glaucoma, but their stability on borderline cases remains limited. Comparatively, GPT-5.3 shows higher consistency and more stable reasoning patterns.

Conclusion · 2 of 2

The application of large language models in ophthalmic diagnostic support still requires cautious evaluation.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Cross-sectional
44.1%
had too much homocysteine in the blood, a far larger share than in dry AMD or healthy eyes
Cohort Study
92.3%
most laser-treated eyes had a favorable eye-structure outcome
Cohort Study
56%
lost three or more lines on the eye-chart vision test
Cohort Study
50%
older patients were more likely to lose three or more lines of vision
Study
81%
trials where mask air leaks were detected blowing toward the eyes
Cohort Study
13.5%
Central retinal artery peak blood-flow speed was 13.5% lower in diabetic patients.
Study
379 +/- 156 microm
lower thickness in the central retina, checked with an eye scan
Case Report
9 of 10 eyes
eyes had better measured vision after treatment