Post

New research · Ophthalmology
JMIR medical informatics · 5d
AI / informaticsJMIR medical informatics · 2026

Mapping the Reliability-Readability Gap in the Education of Patients With Age-Related Macular Degeneration Across 6 Large Language Models: Comparative Evaluation Study.

Zhili Lu, Haixing Cao, Cong Ma … Xiang Ma
Read paper
OphthalmologyAI / informatics

Even the most readable AI models for age-related macular degeneration information are too complex.

Mapping the Reliability-Readability Gap in the Education of Patients With Age-Related Macular Degeneration Across 6 Large Language Models: Comparative Evaluation Study.

Zhili Lu … Xiang Ma
JMIR medical informatics · 2026
Background

Artificial intelligence-generated health information is increasingly used by patients, but its reliability, visible transparency indicators, and readability remain uncertain in specialized ophthalmic conditions such as age-related macular degeneration (AMD).

Purpose

This study aimed to evaluate and compare the informational reliability, visible transparency indicators, overall quality, and readability of responses generated by 6 publicly accessible large language models (large language models) to AMD-related patient-facing prompts under a zero-shot, single-turn prompting scenario.

Methods

Thirty English-language AMD-related prompts were curated from Google Trends, the 2023 Chinese AMD guideline, and the 2025 American Academy of Ophthalmology Preferred Practice Pattern.

n = 180 responses
mean 9.95
Results
mean 9.95
average reading grade level required for the most readable AI information
n = 180 responses
More results

Interrater agreement was substantial to near-perfect across reliability instruments (κ=0.72-0.97).

No model met the recommended sixth-grade readability target.

Significant between-model differences were observed across all reliability and readability metrics (all P<.001).

“
Conclusion · 1 of 3

Under zero-shot, single-turn prompting conditions, the evaluated public large language models showed substantial model-dependent differences in AMD-related patient education quality and readability.

Conclusion · 2 of 3

No model met the sixth-grade readability benchmark, including those with comparatively stronger reliability performance.

Conclusion · 3 of 3

These findings support clinician oversight, readability optimization, and further evaluation before large language models-generated AMD information is used directly in patient-facing settings.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
0·962 for OCT
MerMED-FM was very accurate at diagnosing diseases using eye scans
AI / Informatics
0.98
Artificial intelligence's eyelid-height measurements matched doctors' manual measurements almost perfectly
AI / Informatics
92.1%
combining eye scans and photos correctly told benign from cancerous lesions apart nearly every time
Cohort Study
28.6%
recurred locally in more than 1 in 4 patients over years of follow-up
Cohort Study
-0.395
excision group's post-op eyelid fullness score was lower - greater improvement
Cohort Study
86.7%
of infants probed after 12 months still had unresolved tear duct blockage
Observational
16%
eyelid tissue in rosacea patients showed less of this key repair-signaling protein inside cell nuclei
Cohort Study
25%
about 1 in 4 treated cases had symptoms return after improving