Post

New research · Ophthalmology
JAMA ophthalmology · 22h
AI / informaticsJAMA ophthalmology · 2025

Leveraging Large Language Models to Generate Multiple-Choice Questions for Ophthalmology Education.

Shahrzad Gholami, Daniel B Mummert, Beth Wilson … Karine D Bojikian
Read paper
OphthalmologyAI / informatics

Large language models generate ophthalmology multiple-choice questions comparable to human experts.

Leveraging Large Language Models to Generate Multiple-Choice Questions for Ophthalmology Education.

Shahrzad Gholami et al. · JAMA ophthalmology · 2025
Background

IMPORTANCE: Multiple choice questions (MCQs) are an important and integral component of ophthalmology residency training evaluation and board certification; however, high-quality questions are difficult and time-consuming to draft.

Purpose

To evaluate whether general-domain large language models (LLMs), particularly OpenAI's Generative Pre-trained Transformer 4 (GPT-4), can reliably generate high-quality, novel, and readable MCQs comparable to those of a committee of experienced examination writers.

Methods

DESIGN, SETTING, AND PARTICIPANTS: This survey study, conducted from September 2024 to April 2025, assesses LLM performance in generating MCQs based on the American Academy of Ophthalmology (AAO) Basic and Clinical Science Course (BCSC) compared with a committee of human experts.

difference, 0
Results
AI and human experts had the same median quality scores for eye care questions
More results

The 10 graders had between 1 and 28 years of clinical experience in ophthalmology (median [IQR] experience, 6 years [3-15 years]).

Nearly 95% of LLM-MCQs had similarity scores less than 60, indicating most LLM-MCQs had limited or no resemblance to existing content.

More results

Interrater reliability was moderate (ICC, 0.63; P < .001), and mean (SD) readability scores were similar across sources (37.14 [22.54] vs 42.60 [22.84]; P > .99).

“
Conclusion · 1 of 2

In this survey study, results indicate that an LLM could be used to develop ophthalmology board-style MCQs and expand examination banks to further support ophthalmology residency training.

Conclusion · 2 of 2

Despite most questions having a low similarity score, the quality, novelty, and readability of the LLM-generated questions need to be further assessed.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Cohort Study
median 9
optometry referrals contained more complete documentation for glaucoma
Study
71.27%
medical students' higher accuracy in identifying eye infections after artificial intelligence training
Cross-sectional
64%
most patients reported little difficulty with their eye injections
Cross-sectional
37%
Only 37% of primary care providers and endocrinologists correctly identified diabetic retinopathy status.
Randomized Trial
adjusted difference -4.97
Baduanjin exercise led to a better overall quality of life than routine care
Study
58.26%
Qwen-7B correctly found eye diseases in over half of cases without specific training
Case Report
24 months
the patient's eye cancer remained gone for this period after treatment
Study
Mean 184 vs. 137
Residents in subsidized programs performed more cataract surgeries as primary surgeon.