Post

New research · Ophthalmology
JAMA ophthalmology · 23h
AI / informaticsJAMA ophthalmology · 2025

Ophthalmological Question Answering and Reasoning Using OpenAI o1 vs Other Large Language Models.

Sahana Srinivasan, Xuguang Ai, Minjie Zou … Yih Chung Tham
Read paper
OphthalmologyAI / informatics

OpenAI's o1 large language model achieved 0.877 accuracy on ophthalmology questions.

Ophthalmological Question Answering and Reasoning Using OpenAI o1 vs Other Large Language Models.

Sahana Srinivasan et al. · JAMA ophthalmology · 2025
Background

IMPORTANCE: OpenAI's recent large language model (LLM) o1 has dedicated reasoning capabilities, but it remains untested in specialized medical fields like ophthalmology.

Purpose

To assess the performance and reasoning ability of OpenAI's o1 compared with other LLMs on ophthalmological questions.

Methods

DESIGN, SETTING, AND PARTICIPANTS: In September through October 2024, the LLMs o1, GPT-4o (OpenAI), GPT-4 (OpenAI), GPT-3.5 (OpenAI), Llama 3-8B (Meta), and Gemini 1.5 Pro (Google) were evaluated on 6990 standardized ophthalmology questions from the Medical Multiple-Choice Question Answering (MedMCQA) dataset.

n = 6990 standardized ophthalmology
Results

The o1 large language model correctly answered 87.7% of ophthalmology questions.

accuracy 0.88 (95% CI 0.87 to 0.89)
null = 00.870.89
CI excludes the null - significant
More results

In BERTScore, GPT-4o (Δ = 0.012; 95% CI, 0.012 to 0.013) and GPT-4 (Δ = 0.014; 95% CI, 0.014 to 0.015) outperformed o1 (P < .001).

Similarly, in AlignScore, GPT-4o (Δ = 0.019; 95% CI, 0.016 to 0.021) and GPT-4 (Δ = 0.024; 95% CI, 0.021 to 0.026) again performed better (P < .001).

More results

Conversely, o1 led in BARTScore (mean, -4.787; 95% CI, -4.813 to -4.762; P < .001) and METEOR (mean, 0.221; 95% CI, 0.218 to 0.223; P < .001 except GPT-4o).

“
Conclusion · 1 of 2

This study found that o1 excelled in accuracy but showed inconsistencies in text-generation metrics, trailing GPT-4o and GPT-4; expert reviews found o1's responses to be more clinically useful and better organized than GPT-4o.

Conclusion · 2 of 2

While o1 demonstrated promise, its performance in addressing ophthalmology-specific challenges is not fully optimal, underscoring the potential need for domain-specialized LLMs and targeted evaluations.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
difference, 0
AI and human experts had the same median quality scores for eye care questions
Case-control
0 of 28
distinct structural changes were absent in all 28 healthy control eyes
Cohort Study
RR, 0.91
Black patients were less likely to start medication for their eye condition
AI / Informatics
100% diagnostic sensitivity
ROFI correctly found every case of eye disease
AI / Informatics
10.50-point increase
students trained with digital patients had higher history-taking assessment scores
Cohort Study
HR 4.85
chronic pain conditions mean a nearly five times higher risk of developing dry eye disease
Cohort Study
12 per 10,000
a higher rate of eye infections after dexamethasone implant injections
Guideline
first 5 years
annual eye screening for damage can be delayed for this long if no risk factors