Post

New research · Ophthalmology
JAMA ophthalmology · 22h
AI / informaticsJAMA ophthalmology · 2025

Ophthalmological Question Answering and Reasoning Using OpenAI o1 vs Other Large Language Models.

Sahana Srinivasan, Xuguang Ai, Minjie Zou … Yih Chung Tham
Read paper
OphthalmologyAI / informatics

OpenAI's o1 large language model achieved 0.877 accuracy on ophthalmology questions.

Ophthalmological Question Answering and Reasoning Using OpenAI o1 vs Other Large Language Models.

Sahana Srinivasan et al. · JAMA ophthalmology · 2025
Background

IMPORTANCE: OpenAI's recent large language model (LLM) o1 has dedicated reasoning capabilities, but it remains untested in specialized medical fields like ophthalmology.

Purpose

To assess the performance and reasoning ability of OpenAI's o1 compared with other LLMs on ophthalmological questions.

Methods

DESIGN, SETTING, AND PARTICIPANTS: In September through October 2024, the LLMs o1, GPT-4o (OpenAI), GPT-4 (OpenAI), GPT-3.5 (OpenAI), Llama 3-8B (Meta), and Gemini 1.5 Pro (Google) were evaluated on 6990 standardized ophthalmology questions from the Medical Multiple-Choice Question Answering (MedMCQA) dataset.

n = 6990 standardized ophthalmology
Results

The o1 large language model correctly answered 87.7% of ophthalmology questions.

accuracy 0.88 (95% CI 0.87 to 0.89)
null = 00.870.89
CI excludes the null - significant
More results

In BERTScore, GPT-4o (Δ = 0.012; 95% CI, 0.012 to 0.013) and GPT-4 (Δ = 0.014; 95% CI, 0.014 to 0.015) outperformed o1 (P < .001).

Similarly, in AlignScore, GPT-4o (Δ = 0.019; 95% CI, 0.016 to 0.021) and GPT-4 (Δ = 0.024; 95% CI, 0.021 to 0.026) again performed better (P < .001).

More results

Conversely, o1 led in BARTScore (mean, -4.787; 95% CI, -4.813 to -4.762; P < .001) and METEOR (mean, 0.221; 95% CI, 0.218 to 0.223; P < .001 except GPT-4o).

“
Conclusion · 1 of 2

This study found that o1 excelled in accuracy but showed inconsistencies in text-generation metrics, trailing GPT-4o and GPT-4; expert reviews found o1's responses to be more clinically useful and better organized than GPT-4o.

Conclusion · 2 of 2

While o1 demonstrated promise, its performance in addressing ophthalmology-specific challenges is not fully optimal, underscoring the potential need for domain-specialized LLMs and targeted evaluations.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Cohort Study
median 9
optometry referrals contained more complete documentation for glaucoma
Study
71.27%
medical students' higher accuracy in identifying eye infections after artificial intelligence training
Cross-sectional
64%
most patients reported little difficulty with their eye injections
Cross-sectional
37%
Only 37% of primary care providers and endocrinologists correctly identified diabetic retinopathy status.
Randomized Trial
adjusted difference -4.97
Baduanjin exercise led to a better overall quality of life than routine care
Study
58.26%
Qwen-7B correctly found eye diseases in over half of cases without specific training
Case Report
24 months
the patient's eye cancer remained gone for this period after treatment
Study
Mean 184 vs. 137
Residents in subsidized programs performed more cataract surgeries as primary surgeon.