Post

New research · Ophthalmology
European journal of ophthalmology · 1d
AI / informaticsEuropean journal of ophthalmology · 2026

AI chatbots in strabismus care: A multidomain expert evaluation of caregiver-facing information.

Dhiman Shweta, Dutta Paromita, Thacker Prolima … Mishra Chitaranjan
Read paper
OphthalmologyAI / informatics

ChatGPT rated most accurate and clear among artificial intelligence chatbots answering caregiver questions on strabismus

AI chatbots in strabismus care: A multidomain expert evaluation of caregiver-facing information.

Dhiman Shweta … Mishra Chitaranjan
European journal of ophthalmology · 2026
Purpose

To evaluate and compare the performance of five artificial intelligence (artificial intelligence) chatbots-ChatGPT (OpenAI 4), Google Gemini, Grok (xAI), DeepSeek, and Meta Llama -in delivering accurate, clear, educational, and safe responses to caregiver-facing queries related to strabismus.

Methods

Sixteen standardized caregiver questions on strabismus were presented to each chatbot in independent sessions.

Results

of 16 questions, most experts rated ChatGPT's answers highly accurate

65%
ChatGPT
41%
Llama
More results

For Educational Value, Llama (43.8%) and Gemini (42.5%) performed slightly better, while Safety ratings were highest for Gemini (40%) and Llama (37.5%).

Cumulative link mixed models analysis showed significant between-chatbot differences for Accuracy, Clarity, and Educational Value ( p < 0.05) but not for Safety.

More results

Compared with ChatGPT, lower odds of higher ratings were seen for Grok (OR 0.48) and DeepSeek (OR 0.61).

Inter-rater reliability indicated moderate agreement (Fleiss' κ = 0.59) and strong consensus (Gwet's AC1 = 0.87).

“
Conclusion · 1 of 2

ChatGPT showed superior accuracy and clarity, while Gemini and Llama excelled in educational value and safety.

Conclusion · 2 of 2

High expert agreement supports artificial intelligence chatbots as adjuncts in pediatric ophthalmology education requiring continued validation.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Cross-sectional
44.1%
had too much homocysteine in the blood, a far larger share than in dry AMD or healthy eyes
Cohort Study
92.3%
most laser-treated eyes had a favorable eye-structure outcome
Cohort Study
56%
lost three or more lines on the eye-chart vision test
Cohort Study
50%
older patients were more likely to lose three or more lines of vision
Study
81%
trials where mask air leaks were detected blowing toward the eyes
Cohort Study
13.5%
Central retinal artery peak blood-flow speed was 13.5% lower in diabetic patients.
Study
379 +/- 156 microm
lower thickness in the central retina, checked with an eye scan
Case Report
9 of 10 eyes
eyes had better measured vision after treatment