Leveraging Large Language Models to Generate Multiple-Choice Questions for Ophthalmology Education.
Large language models generate ophthalmology multiple-choice questions comparable to human experts.
Leveraging Large Language Models to Generate Multiple-Choice Questions for Ophthalmology Education.
IMPORTANCE: Multiple choice questions (multiple choice questions) are an important and integral component of ophthalmology residency training evaluation and board certification; however, high-quality questions are difficult and time-consuming to draft.
To evaluate whether general-domain large language models (large language models), particularly OpenAI's Generative Pre-trained Transformer 4 (GPT-4), can reliably generate high-quality, novel, and readable multiple choice questions comparable to those of a committee of experienced examination writers.
DESIGN, SETTING, AND PARTICIPANTS: This survey study, conducted from September 2024 to April 2025, assesses large language models performance in generating multiple choice questions based on the American Academy of Ophthalmology (Academy of Ophthalmology) Basic and Clinical Science Course (Basic and Clinical Science Course) compared with a committee of human experts.
The 10 graders had between 1 and 28 years of clinical experience in ophthalmology (median [IQR] experience, 6 years [3-15 years]).
Nearly 95% of LLM-MCQs had similarity scores less than 60, indicating most LLM-MCQs had limited or no resemblance to existing content.
Interrater reliability was moderate (intraclass correlation coefficient, 0.63; P < .001), and mean (SD) readability scores were similar across sources (37.14 [22.54] vs 42.60 [22.84]; P > .99).
In this survey study, results indicate that an large language models could be used to develop ophthalmology board-style multiple choice questions and expand examination banks to further support ophthalmology residency training.
Despite most questions having a low similarity score, the quality, novelty, and readability of the large language models-generated questions need to be further assessed.