Mapping the Reliability-Readability Gap in the Education of Patients With Age-Related Macular Degeneration Across 6 Large Language Models: Comparative Evaluation Study.
Even the most readable AI models for age-related macular degeneration information are too complex.
Mapping the Reliability-Readability Gap in the Education of Patients With Age-Related Macular Degeneration Across 6 Large Language Models: Comparative Evaluation Study.
Artificial intelligence-generated health information is increasingly used by patients, but its reliability, visible transparency indicators, and readability remain uncertain in specialized ophthalmic conditions such as age-related macular degeneration (AMD).
This study aimed to evaluate and compare the informational reliability, visible transparency indicators, overall quality, and readability of responses generated by 6 publicly accessible large language models (large language models) to AMD-related patient-facing prompts under a zero-shot, single-turn prompting scenario.
Thirty English-language AMD-related prompts were curated from Google Trends, the 2023 Chinese AMD guideline, and the 2025 American Academy of Ophthalmology Preferred Practice Pattern.
Interrater agreement was substantial to near-perfect across reliability instruments (κ=0.72-0.97).
No model met the recommended sixth-grade readability target.
Significant between-model differences were observed across all reliability and readability metrics (all P<.001).
Under zero-shot, single-turn prompting conditions, the evaluated public large language models showed substantial model-dependent differences in AMD-related patient education quality and readability.
No model met the sixth-grade readability benchmark, including those with comparatively stronger reliability performance.
These findings support clinician oversight, readability optimization, and further evaluation before large language models-generated AMD information is used directly in patient-facing settings.