Post

New research · Ophthalmology
Arquivos brasileiros de oftalmologia · 1d
AI / informaticsArquivos brasileiros de oftalmologia · 2026

Assessing a large language model for glaucoma knowledge: ChatGPT-5 versus residents.

Mauro Gobira, Rodrigo Moreira, Flavio J L Galhardo Carvalho Filho … Ivan M Tavares
Read paper
OphthalmologyAI / informatics

ChatGPT-5 had higher odds of correctly answering glaucoma questions than ophthalmology residents.

Assessing a large language model for glaucoma knowledge: ChatGPT-5 versus residents.

Mauro Gobira et al. · Arquivos brasileiros de oftalmologia · 2026
Purpose

To assess the performance of a contemporary large language model (ChatGPT-5) against ophthalmology residents on a standardized set of glaucoma multiple-choice questions.

Methods

We conducted a cross-sectional comparative study with 189 text-only glaucoma multiple-choice questions from the Cybersight question bank.

Results
ChatGPT-5 was more likely to correctly answer glaucoma questions than eye doctor trainees
More results

ChatGPT-5 received 164 of 189 correct responses (86.8%; 95% CI, 81.2-90.9).

Residents' overall accuracy was 62.9% (713/1,134; 95% CI, 60.0-65.6).

The top-performing resident earned 76.7%.

More results

ChatGPT-5 correctly answered 17/189 items (9.0%), but fewer than half of residents were correct ("large language model-only wins"), whereas residents were more successful on items that ChatGPT-5 overlooked.

“
Conclusion · 1 of 3

ChatGPT-5 outperformed ophthalmology residents on text-based glaucoma multiple-choice questions, indicating its potential as a subspecialty education and assessment tool.

Conclusion · 2 of 3

Generalizability is limited by the single question bank, text-only items, a small resident cohort, and the evaluation of one large language model version at a single time point.

Conclusion · 3 of 3

Before incorporating these findings into clinical decision-making, larger, multimodal, and longitudinal studies are required.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
difference, 0
AI and human experts had the same median quality scores for eye care questions
Case-control
0 of 28
distinct structural changes were absent in all 28 healthy control eyes
Cohort Study
RR, 0.91
Black patients were less likely to start medication for their eye condition
AI / Informatics
100% diagnostic sensitivity
ROFI correctly found every case of eye disease
AI / Informatics
10.50-point increase
students trained with digital patients had higher history-taking assessment scores
AI / Informatics
0.877
The o1 large language model correctly answered 87.7% of ophthalmology questions.
Cohort Study
HR 4.85
chronic pain conditions mean a nearly five times higher risk of developing dry eye disease
Cohort Study
12 per 10,000
a higher rate of eye infections after dexamethasone implant injections