Post

New research · Ophthalmology
European journal of ophthalmology · 23h
AI / informaticsEuropean journal of ophthalmology · 2026

AI chatbots in strabismus care: A multidomain expert evaluation of caregiver-facing information.

Dhiman Shweta, Dutta Paromita, Thacker Prolima … Mishra Chitaranjan
Read paper
OphthalmologyAI / informatics

ChatGPT rated most accurate and clear among artificial intelligence chatbots answering caregiver questions on strabismus

AI chatbots in strabismus care: A multidomain expert evaluation of caregiver-facing information.

Dhiman Shweta … Mishra Chitaranjan
European journal of ophthalmology · 2026
Purpose

To evaluate and compare the performance of five artificial intelligence (artificial intelligence) chatbots-ChatGPT (OpenAI 4), Google Gemini, Grok (xAI), DeepSeek, and Meta Llama -in delivering accurate, clear, educational, and safe responses to caregiver-facing queries related to strabismus.

Methods

Sixteen standardized caregiver questions on strabismus were presented to each chatbot in independent sessions.

Results

of 16 questions, most experts rated ChatGPT's answers highly accurate

65%
ChatGPT
41%
Llama
More results

For Educational Value, Llama (43.8%) and Gemini (42.5%) performed slightly better, while Safety ratings were highest for Gemini (40%) and Llama (37.5%).

Cumulative link mixed models analysis showed significant between-chatbot differences for Accuracy, Clarity, and Educational Value ( p < 0.05) but not for Safety.

More results

Compared with ChatGPT, lower odds of higher ratings were seen for Grok (OR 0.48) and DeepSeek (OR 0.61).

Inter-rater reliability indicated moderate agreement (Fleiss' κ = 0.59) and strong consensus (Gwet's AC1 = 0.87).

“
Conclusion · 1 of 2

ChatGPT showed superior accuracy and clarity, while Gemini and Llama excelled in educational value and safety.

Conclusion · 2 of 2

High expert agreement supports artificial intelligence chatbots as adjuncts in pediatric ophthalmology education requiring continued validation.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
0·962 for OCT
MerMED-FM was very accurate at diagnosing diseases using eye scans
AI / Informatics
0.98
Artificial intelligence's eyelid-height measurements matched doctors' manual measurements almost perfectly
AI / Informatics
92.1%
combining eye scans and photos correctly told benign from cancerous lesions apart nearly every time
Cohort Study
28.6%
recurred locally in more than 1 in 4 patients over years of follow-up
Cohort Study
-0.395
excision group's post-op eyelid fullness score was lower - greater improvement
Cohort Study
86.7%
of infants probed after 12 months still had unresolved tear duct blockage
Observational
16%
eyelid tissue in rosacea patients showed less of this key repair-signaling protein inside cell nuclei
Cohort Study
25%
about 1 in 4 treated cases had symptoms return after improving