Post

New research · Ophthalmology
JAMA ophthalmology · 23h
AI / informaticsJAMA ophthalmology · 2026

Performance of Foundation Models vs Physicians in Textual and Multimodal Ophthalmological Questions.

Henry Rocha, Yu Jeat Chong, Arun James Thirunavukarasu … Darren Shu Jeng Ting
Read paper
OphthalmologyAI / informatics

Foundation Model Claude 3.5 Sonnet showed 35.2% higher accuracy than junior physicians.

Performance of Foundation Models vs Physicians in Textual and Multimodal Ophthalmological Questions.

Henry Rocha et al. · JAMA ophthalmology · 2026
Background

IMPORTANCE: There is an increasing amount of literature evaluating the clinical knowledge and reasoning performance of large language models (LLMs) in ophthalmology, but to date, investigations into its multimodal abilities clinically-such as interpreting images and tables-have been limited.

Purpose

To evaluate the multimodal performance of the following 7 foundation models (FMs): GPT-4o (OpenAI), Gemini 1.5 Pro (Google), Claude 3.5 Sonnet (Anthropic), Llama-3.2-11B (Meta), DeepSeek V3 (High-Flyer), Qwen2.5-Max (Alibaba Cloud), and Qwen2.5-VL-72B (Alibaba Cloud) in answering offline Fellowship of the Royal College of Ophthalmologists part 2 written multiple-choice textual and multimodal questions, with head-to-head comparisons with physicians.

Methods

DESIGN, SETTING, AND PARTICIPANTS: This cross-sectional study was conducted between September 2024 and March 2025 using questions sourced from a textbook used as an examination preparation resource for the Fellowship of the Royal College of Ophthalmologists part 2 written examination.

Results

Claude 3.5 Sonnet was more accurate than junior doctors on textual questions

difference 9 (95% CI 2.40 to 15.60)
null 0
2.40
15.60
CI excludes the null - significant
More results

GPT-4o (accuracy, 69.9%) outperformed GPT-4 (OpenAI; difference, 8.5%; 95% CI, 1.1%-15.8%; P = .02) and GPT-3.5 (OpenAI; difference, 21.8%; 95% CI, 14.3%-29.2%; P < .001).

“
Conclusion · 1 of 2

Results of this cross-sectional study suggest that for textual questions, current FMs exhibited notable improvements in ophthalmological knowledge reasoning when compared with older LLMs and ophthalmology trainees, with performance comparable with that of expert ophthalmologists.

Conclusion · 2 of 2

These models demonstrated potential for medical assistance for answering ophthalmological textual queries, but their multimodal abilities remain limited.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
difference, 0
AI and human experts had the same median quality scores for eye care questions
Case-control
0 of 28
distinct structural changes were absent in all 28 healthy control eyes
Cohort Study
RR, 0.91
Black patients were less likely to start medication for their eye condition
AI / Informatics
100% diagnostic sensitivity
ROFI correctly found every case of eye disease
AI / Informatics
10.50-point increase
students trained with digital patients had higher history-taking assessment scores
AI / Informatics
0.877
The o1 large language model correctly answered 87.7% of ophthalmology questions.
Cohort Study
HR 4.85
chronic pain conditions mean a nearly five times higher risk of developing dry eye disease
Cohort Study
12 per 10,000
a higher rate of eye infections after dexamethasone implant injections