DeepSeek-R1 vs OpenAI o1 for Ophthalmic Diagnoses and Management Plans.
DeepSeek-R1 had higher diagnostic accuracy than OpenAI o1 in ophthalmic cases.
DeepSeek-R1 vs OpenAI o1 for Ophthalmic Diagnoses and Management Plans.
IMPORTANCE: Large language models (large language models) are increasingly being explored in clinical decision-making, but few studies have evaluated their performance on complex ophthalmology cases from clinical practice settings.
To evaluate the diagnostic accuracy, management decision-making, and cost of DeepSeek-R1 vs OpenAI o1 across diverse ophthalmic subspecialties.
DESIGN, SETTING, AND PARTICIPANTS: This was a cross-sectional evaluation conducted using standardized prompts and model configurations.
DeepSeek-R1 correctly identified the eye condition in more cases than OpenAI o1
For next-step decisions, DeepSeek-R1 was correct in 82.7% of cases (349 of 422 cases) vs OpenAI o1's accuracy of 75.8% (320 of 422 cases), a 6.9% difference (95% CI, 1.4%-12.3%; P = .01).
Intermodel agreement was moderate (κ = 0.422; 95% CI, 0.375-0.469; P < .001).
DeepSeek-R1 offered lower costs per query than OpenAI o1, with savings exceeding 66-fold (up to 98.5%) during off-peak pricing.
DeepSeek-R1 outperformed OpenAI o1 in diagnosis and management across subspecialties while lowering operating costs, supporting the potential of open-weight, reinforcement learning-augmented large language models as scalable and cost-saving tools for clinical decision support.
Further investigations should evaluate safety guardrails and assess performance of self-hosted adaptations of DeepSeek-R1 with domain-specific ophthalmic expertise to optimize clinical utility.