Clinical Safety and Reliability of Large Language Models in Answering Hemorrhoid-Related Patient Questions: A Comparative Study of ChatGPT, Gemini, and DeepSeek.
ChatGPT scored highest among large language models answering hemorrhoid patient questions.
Clinical Safety and Reliability of Large Language Models in Answering Hemorrhoid-Related Patient Questions: A Comparative Study of ChatGPT, Gemini, and DeepSeek.
Large language models (large language models) are increasingly used by patients for obtaining medical information; however, concerns remain regarding their clinical safety, reliability, and appropriateness of patient guidance.
To compare the clinical accuracy, safety, and overall clinical adequacy of responses generated by ChatGPT, Gemini, and DeepSeek to hemorrhoid-related patient questions.
In this cross-sectional comparative study, 25 hemorrhoid-related patient questions were developed and categorized into three predefined subgroups: basic informational questions, clinically significant scenarios, and misleading/risky patient statements.
chatgpt gave the most reliable, consistently accurate answers
Overall response quality differed significantly among models (χ 2 (2) = 29.119, p < 0.001, Kendall's W = 0.582).
Significant differences were primarily observed between DeepSeek and the other models, whereas ChatGPT and Gemini showed comparable performance.
Model divergence became more pronounced in clinically significant scenarios involving alarm symptoms, rectal bleeding, persistent symptoms, and acute anorectal pain.
However, qualitative assessment demonstrated differences in communication style and risk communication.
Large language models demonstrated generally high clinical accuracy in answering hemorrhoid-related patient questions; however, notable model-specific differences were observed in clinical guidance, communication style, and risk communication, particularly in high-risk clinical scenarios.
These findings suggest that large language models may serve as useful supportive tools for patient education and health information delivery, although they should currently be regarded as systems that support rather than replace human clinical judgment.