Post

New research · Ophthalmology
medRxiv : the preprint server for health sciences · 23h
Cohort studymedRxiv : the preprint server for health sciences · 2026

Extraction of Glaucoma Diagnosis, Type, and Severity from Clinical Notes using Secure Cloud-based Large Language Models.

Gustavo A Samico, Nicholas Solages, Rafael Scherer … Swarup S Swaminathan
Read paper
OphthalmologyCohort study

Claude large language model achieved 97.5% accuracy in glaucoma diagnosis.

Extraction of Glaucoma Diagnosis, Type, and Severity from Clinical Notes using Secure Cloud-based Large Language Models.

Gustavo A Samico et al. · medRxiv : the preprint server for health sciences · 2026
Purpose

To evaluate the performance of secure cloud-based large language models (LLMs) in extracting glaucoma diagnosis, type, and severity from free-text clinical notes in the electronic health record (EHR).

Methods

1,250 subjects from the Bascom Palmer Ophthalmic Repository.

n = 1,250 subjects
Results

Claude correctly identified glaucoma in a very high percentage of patients

Gwet AC1 0.93 (95% CI 0.92 to 0.94)
null 0
0.92
0.94
CI excludes the null - significant
More results

F1 scores for glaucoma detection ranged from 95.4% to 98.9% across models.

For glaucoma type classification, accuracies were 97.1%, 94.2%, 94.2%, 94.0%, and 94.4% for Claude, DeepSeek, GPT, Grok, and Qwen, respectively.

F1 scores for the most prevalent type (POAG) ranged from 96.3% to 98.9%.

More results

ICD-10 codes demonstrated substantially lower performance for type and severity staging, with overall accuracies of 89.2% and 58.5%, respectively.

“
Conclusion · 1 of 2

Secure cloud-based LLMs accurately extracted glaucoma diagnosis, type, and severity information from free-text ophthalmology notes, achieving performance approaching expert clinician adjudication while substantially outperforming ICD-based phenotyping approaches, particularly for disease severity classification.

Conclusion · 2 of 2

These findings demonstrate the potential of LLMs to transform unstructured clinical documentation into scalable, research-ready phenotypic data for large-scale glaucoma cohort development and EHR-based ophthalmic research.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
difference, 0
AI and human experts had the same median quality scores for eye care questions
Case-control
0 of 28
distinct structural changes were absent in all 28 healthy control eyes
Cohort Study
RR, 0.91
Black patients were less likely to start medication for their eye condition
AI / Informatics
100% diagnostic sensitivity
ROFI correctly found every case of eye disease
AI / Informatics
10.50-point increase
students trained with digital patients had higher history-taking assessment scores
AI / Informatics
0.877
The o1 large language model correctly answered 87.7% of ophthalmology questions.
Cohort Study
HR 4.85
chronic pain conditions mean a nearly five times higher risk of developing dry eye disease
Cohort Study
12 per 10,000
a higher rate of eye infections after dexamethasone implant injections