Post

New research · Ophthalmology
JAMA ophthalmology · 23h
AI / informaticsJAMA ophthalmology · 2025

Leveraging Large Language Models to Generate Multiple-Choice Questions for Ophthalmology Education.

Shahrzad Gholami, Daniel B Mummert, Beth Wilson … Karine D Bojikian
Read paper
OphthalmologyAI / informatics

Large language models generate ophthalmology multiple-choice questions comparable to human experts.

Leveraging Large Language Models to Generate Multiple-Choice Questions for Ophthalmology Education.

Shahrzad Gholami et al. · JAMA ophthalmology · 2025
Background

IMPORTANCE: Multiple choice questions (MCQs) are an important and integral component of ophthalmology residency training evaluation and board certification; however, high-quality questions are difficult and time-consuming to draft.

Purpose

To evaluate whether general-domain large language models (LLMs), particularly OpenAI's Generative Pre-trained Transformer 4 (GPT-4), can reliably generate high-quality, novel, and readable MCQs comparable to those of a committee of experienced examination writers.

Methods

DESIGN, SETTING, AND PARTICIPANTS: This survey study, conducted from September 2024 to April 2025, assesses LLM performance in generating MCQs based on the American Academy of Ophthalmology (AAO) Basic and Clinical Science Course (BCSC) compared with a committee of human experts.

difference, 0
Results
AI and human experts had the same median quality scores for eye care questions
More results

The 10 graders had between 1 and 28 years of clinical experience in ophthalmology (median [IQR] experience, 6 years [3-15 years]).

Nearly 95% of LLM-MCQs had similarity scores less than 60, indicating most LLM-MCQs had limited or no resemblance to existing content.

More results

Interrater reliability was moderate (ICC, 0.63; P < .001), and mean (SD) readability scores were similar across sources (37.14 [22.54] vs 42.60 [22.84]; P > .99).

“
Conclusion · 1 of 2

In this survey study, results indicate that an LLM could be used to develop ophthalmology board-style MCQs and expand examination banks to further support ophthalmology residency training.

Conclusion · 2 of 2

Despite most questions having a low similarity score, the quality, novelty, and readability of the LLM-generated questions need to be further assessed.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
10.50-point increase
students trained with digital patients had higher history-taking assessment scores
Case-control
0 of 28
distinct structural changes were absent in all 28 healthy control eyes
Cohort Study
RR, 0.91
Black patients were less likely to start medication for their eye condition
AI / Informatics
100% diagnostic sensitivity
ROFI correctly found every case of eye disease
AI / Informatics
0.877
The o1 large language model correctly answered 87.7% of ophthalmology questions.
Cohort Study
HR 4.85
chronic pain conditions mean a nearly five times higher risk of developing dry eye disease
Cohort Study
12 per 10,000
a higher rate of eye infections after dexamethasone implant injections
Guideline
first 5 years
annual eye screening for damage can be delayed for this long if no risk factors