Large language model assistance improved resident physician accuracy on text-based ophthalmology exams.
Large Language Models for Ophthalmology Training in China: A Prospective Evaluation.
Zuhui Zhang et al. · Ophthalmology science · 2026
Purpose
This study explored large language models (LLMs) as a scalable solution to the global shortage and uneven distribution of ophthalmologists, particularly their actual effectiveness and potential risks in ophthalmic training.
Methods
Phase 1: all LLMs were tested on the Chinese and English versions of the Chinese National Health Professional Technical Qualification Examination (Intermediate Level) in Ophthalmology (CNHPTQE-O).
Results
resident physicians' accuracy on text-based eye exams improved with AI assistance
More results
Several Chinese LLMs, especially ERNIE Bot 4.5 Turbo, demonstrated superior performance on the CNHPTQE-O, achieving accuracies of 98.00% (Chinese) and 86.50% (English).
ERNIE Bot 4.5 Turbo significantly outperformed all RPs on the Chinese examination ( P = 0.001).
Questionnaire feedback was positive.
“
Conclusion · 1 of 2
Large language models possess a solid foundation in ophthalmic knowledge and can effectively enhance trainee performance in text-based assessments, demonstrating clear potential as a training aid.
Conclusion · 2 of 2
However, their limitations in image-assisted diagnostic tasks and the associated risk of "artificial ignorance" should not be overlooked.