Post

New research · Ophthalmology
medRxiv : the preprint server for health sciences · 23h
Cohort studymedRxiv : the preprint server for health sciences · 2026

Extraction of Glaucoma Diagnosis, Type, and Severity from Clinical Notes using Secure Cloud-based Large Language Models.

Gustavo A Samico, Nicholas Solages, Rafael Scherer … Swarup S Swaminathan
Read paper
OphthalmologyCohort study

Claude large language model achieved 97.5% accuracy in glaucoma diagnosis.

Extraction of Glaucoma Diagnosis, Type, and Severity from Clinical Notes using Secure Cloud-based Large Language Models.

Gustavo A Samico et al. · medRxiv : the preprint server for health sciences · 2026
Purpose

To evaluate the performance of secure cloud-based large language models (LLMs) in extracting glaucoma diagnosis, type, and severity from free-text clinical notes in the electronic health record (EHR).

Methods

1,250 subjects from the Bascom Palmer Ophthalmic Repository.

n = 1,250 subjects
Results

Claude correctly identified glaucoma in a very high percentage of patients

Gwet AC1 0.93 (95% CI 0.92 to 0.94)
null 0
0.92
0.94
CI excludes the null - significant
More results

F1 scores for glaucoma detection ranged from 95.4% to 98.9% across models.

For glaucoma type classification, accuracies were 97.1%, 94.2%, 94.2%, 94.0%, and 94.4% for Claude, DeepSeek, GPT, Grok, and Qwen, respectively.

F1 scores for the most prevalent type (POAG) ranged from 96.3% to 98.9%.

More results

ICD-10 codes demonstrated substantially lower performance for type and severity staging, with overall accuracies of 89.2% and 58.5%, respectively.

“
Conclusion · 1 of 2

Secure cloud-based LLMs accurately extracted glaucoma diagnosis, type, and severity information from free-text ophthalmology notes, achieving performance approaching expert clinician adjudication while substantially outperforming ICD-based phenotyping approaches, particularly for disease severity classification.

Conclusion · 2 of 2

These findings demonstrate the potential of LLMs to transform unstructured clinical documentation into scalable, research-ready phenotypic data for large-scale glaucoma cohort development and EHR-based ophthalmic research.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Cohort Study
median 9
optometry referrals contained more complete documentation for glaucoma
Study
71.27%
medical students' higher accuracy in identifying eye infections after artificial intelligence training
Cross-sectional
64%
most patients reported little difficulty with their eye injections
Cross-sectional
37%
Only 37% of primary care providers and endocrinologists correctly identified diabetic retinopathy status.
Randomized Trial
adjusted difference -4.97
Baduanjin exercise led to a better overall quality of life than routine care
Study
58.26%
Qwen-7B correctly found eye diseases in over half of cases without specific training
Case Report
24 months
the patient's eye cancer remained gone for this period after treatment
Study
Mean 184 vs. 137
Residents in subsidized programs performed more cataract surgeries as primary surgeon.