Leveraging Large Language Models to Generate Multiple-Choice Questions for Ophthalmology Education.
Large language models generate ophthalmology multiple-choice questions comparable to human experts.
Leveraging Large Language Models to Generate Multiple-Choice Questions for Ophthalmology Education.
IMPORTANCE: Multiple choice questions (MCQs) are an important and integral component of ophthalmology residency training evaluation and board certification; however, high-quality questions are difficult and time-consuming to draft.
To evaluate whether general-domain large language models (LLMs), particularly OpenAI's Generative Pre-trained Transformer 4 (GPT-4), can reliably generate high-quality, novel, and readable MCQs comparable to those of a committee of experienced examination writers.
DESIGN, SETTING, AND PARTICIPANTS: This survey study, conducted from September 2024 to April 2025, assesses LLM performance in generating MCQs based on the American Academy of Ophthalmology (AAO) Basic and Clinical Science Course (BCSC) compared with a committee of human experts.
The 10 graders had between 1 and 28 years of clinical experience in ophthalmology (median [IQR] experience, 6 years [3-15 years]).
Nearly 95% of LLM-MCQs had similarity scores less than 60, indicating most LLM-MCQs had limited or no resemblance to existing content.
Interrater reliability was moderate (ICC, 0.63; P < .001), and mean (SD) readability scores were similar across sources (37.14 [22.54] vs 42.60 [22.84]; P > .99).
In this survey study, results indicate that an LLM could be used to develop ophthalmology board-style MCQs and expand examination banks to further support ophthalmology residency training.
Despite most questions having a low similarity score, the quality, novelty, and readability of the LLM-generated questions need to be further assessed.