Evaluation of Five Large Language Models for Parental Education in Pediatric Anesthesia: Reliability and Readability Study.
All five large language models answered parental pediatric anesthesia questions with over 90% clinical accuracy.
Evaluation of Five Large Language Models for Parental Education in Pediatric Anesthesia: Reliability and Readability Study.
Although large language models (large language models) show potential for patient education, their accuracy, usability, and comprehensibility lack validation in high-risk pediatric anesthesia.
This study aims to evaluate the accuracy, reliability, and readability of responses generated by 5 large language models to parental inquiries regarding pediatric anesthesia, and to assess their suitability for clinical use in perioperative caregiver education.
Two expert anesthesiologists identified 33 parental questions on pediatric anesthesia by screening authoritative resources and Google Trends.
More than 9 in 10 chatbot answers were rated clinically accurate by expert anesthesiologists.
Excluding Gemini and Copilot, the remaining 3 models (ChatGPT, DeepSeek, and Perplexity) each produced unsafe content in 3.03% (n=1) of the 33 queries.
Hallucinations were detected in all models except Gemini, with DeepSeek and Perplexity showing the highest hallucination rate (3/33, 9.09%).
Furthermore, Perplexity showed superior reliability on DISCERN (median 41; P<.05), yet no model achieved a "good" rating.
Gemini achieved the highest Ensuring Quality Information for Patients (median 66.67%; P<.05) despite lower Global Quality Score (median 3).
In this study, 5 large language models generally provided clinically accurate information when responding to parental questions about pediatric anesthesia.
However, limitations were also identified, including hallucinated content, safety-related deficiencies, limited source transparency, and readability levels exceeding recommended standards.
Therefore, large language models-generated information should be interpreted with caution and should not replace clinician guidance.