Post

New research · Ophthalmology
JAMA ophthalmology · 23h
AI / informaticsJAMA ophthalmology · 2026

Performance of Foundation Models vs Physicians in Textual and Multimodal Ophthalmological Questions.

Henry Rocha, Yu Jeat Chong, Arun James Thirunavukarasu … Darren Shu Jeng Ting
Read paper
OphthalmologyAI / informatics

Foundation Model Claude 3.5 Sonnet showed 35.2% higher accuracy than junior physicians.

Performance of Foundation Models vs Physicians in Textual and Multimodal Ophthalmological Questions.

Henry Rocha et al. · JAMA ophthalmology · 2026
Background

IMPORTANCE: There is an increasing amount of literature evaluating the clinical knowledge and reasoning performance of large language models (LLMs) in ophthalmology, but to date, investigations into its multimodal abilities clinically-such as interpreting images and tables-have been limited.

Purpose

To evaluate the multimodal performance of the following 7 foundation models (FMs): GPT-4o (OpenAI), Gemini 1.5 Pro (Google), Claude 3.5 Sonnet (Anthropic), Llama-3.2-11B (Meta), DeepSeek V3 (High-Flyer), Qwen2.5-Max (Alibaba Cloud), and Qwen2.5-VL-72B (Alibaba Cloud) in answering offline Fellowship of the Royal College of Ophthalmologists part 2 written multiple-choice textual and multimodal questions, with head-to-head comparisons with physicians.

Methods

DESIGN, SETTING, AND PARTICIPANTS: This cross-sectional study was conducted between September 2024 and March 2025 using questions sourced from a textbook used as an examination preparation resource for the Fellowship of the Royal College of Ophthalmologists part 2 written examination.

Results

Claude 3.5 Sonnet was more accurate than junior doctors on textual questions

difference 9 (95% CI 2.40 to 15.60)
null 0
2.40
15.60
CI excludes the null - significant
More results

GPT-4o (accuracy, 69.9%) outperformed GPT-4 (OpenAI; difference, 8.5%; 95% CI, 1.1%-15.8%; P = .02) and GPT-3.5 (OpenAI; difference, 21.8%; 95% CI, 14.3%-29.2%; P < .001).

“
Conclusion · 1 of 2

Results of this cross-sectional study suggest that for textual questions, current FMs exhibited notable improvements in ophthalmological knowledge reasoning when compared with older LLMs and ophthalmology trainees, with performance comparable with that of expert ophthalmologists.

Conclusion · 2 of 2

These models demonstrated potential for medical assistance for answering ophthalmological textual queries, but their multimodal abilities remain limited.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Cohort Study
median 9
optometry referrals contained more complete documentation for glaucoma
Study
71.27%
medical students' higher accuracy in identifying eye infections after artificial intelligence training
Cross-sectional
64%
most patients reported little difficulty with their eye injections
Cross-sectional
37%
Only 37% of primary care providers and endocrinologists correctly identified diabetic retinopathy status.
Randomized Trial
adjusted difference -4.97
Baduanjin exercise led to a better overall quality of life than routine care
Study
58.26%
Qwen-7B correctly found eye diseases in over half of cases without specific training
Case Report
24 months
the patient's eye cancer remained gone for this period after treatment
Study
Mean 184 vs. 137
Residents in subsidized programs performed more cataract surgeries as primary surgeon.