Post

New research · Ophthalmology
JMIR medical informatics · 5d
AI / informaticsJMIR medical informatics · 2026

Mapping the Reliability-Readability Gap in the Education of Patients With Age-Related Macular Degeneration Across 6 Large Language Models: Comparative Evaluation Study.

Zhili Lu, Haixing Cao, Cong Ma … Xiang Ma
Read paper
OphthalmologyAI / informatics

Even the most readable AI models for age-related macular degeneration information are too complex.

Mapping the Reliability-Readability Gap in the Education of Patients With Age-Related Macular Degeneration Across 6 Large Language Models: Comparative Evaluation Study.

Zhili Lu … Xiang Ma
JMIR medical informatics · 2026
Background

Artificial intelligence-generated health information is increasingly used by patients, but its reliability, visible transparency indicators, and readability remain uncertain in specialized ophthalmic conditions such as age-related macular degeneration (AMD).

Purpose

This study aimed to evaluate and compare the informational reliability, visible transparency indicators, overall quality, and readability of responses generated by 6 publicly accessible large language models (large language models) to AMD-related patient-facing prompts under a zero-shot, single-turn prompting scenario.

Methods

Thirty English-language AMD-related prompts were curated from Google Trends, the 2023 Chinese AMD guideline, and the 2025 American Academy of Ophthalmology Preferred Practice Pattern.

n = 180 responses
mean 9.95
Results
mean 9.95
average reading grade level required for the most readable AI information
n = 180 responses
More results

Interrater agreement was substantial to near-perfect across reliability instruments (κ=0.72-0.97).

No model met the recommended sixth-grade readability target.

Significant between-model differences were observed across all reliability and readability metrics (all P<.001).

“
Conclusion · 1 of 3

Under zero-shot, single-turn prompting conditions, the evaluated public large language models showed substantial model-dependent differences in AMD-related patient education quality and readability.

Conclusion · 2 of 3

No model met the sixth-grade readability benchmark, including those with comparatively stronger reliability performance.

Conclusion · 3 of 3

These findings support clinician oversight, readability optimization, and further evaluation before large language models-generated AMD information is used directly in patient-facing settings.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Cross-sectional
44.1%
had too much homocysteine in the blood, a far larger share than in dry AMD or healthy eyes
Cohort Study
92.3%
most laser-treated eyes had a favorable eye-structure outcome
Cohort Study
56%
lost three or more lines on the eye-chart vision test
Cohort Study
50%
older patients were more likely to lose three or more lines of vision
Study
81%
trials where mask air leaks were detected blowing toward the eyes
Cohort Study
13.5%
Central retinal artery peak blood-flow speed was 13.5% lower in diabetic patients.
Study
379 +/- 156 microm
lower thickness in the central retina, checked with an eye scan
Case Report
9 of 10 eyes
eyes had better measured vision after treatment