Post

New research · Ophthalmology
Asia-Pacific journal of ophthalmology (Philadelphia, Pa.) · 5d
AI / informaticsAsia-Pacific journal of ophthalmology (Philadelphia, Pa.) · 2026

Benchmarking and fine-tuning vision-language models on a visual question answering dataset for myopic maculopathy.

Tsun Hei Yip, Pusheng Xu, Zirong Liu … Xiu Juan Zhang
Read paper
OphthalmologyAI / informatics

Fine-tuned InternVL3-8B model matched top proprietary vision-language models in myopic maculopathy visual question answering accuracy

Benchmarking and fine-tuning vision-language models on a visual question answering dataset for myopic maculopathy.

Tsun Hei Yip … Xiu Juan Zhang
Asia-Pacific journal of ophthalmology (Philadelphia, Pa.) · 2026
Purpose

To build a visual question answering (visual question answering) dataset for fine-tuning and evaluating vision-language models (vision-language models) in myopic maculopathy (myopic maculopathy).

Methods

Cross-sectional study.

n = 2591 CFPs and 19,648
Results

fine-tuned model answered about three in four eye-scan questions correctly

0.746accura
Fine-tuned InternVL3-8
0.724accura
Gemini 3 Pro
0.596accura
Claude Sonnet 4.5
0.566accura
Qwen3-VL-30B-A3B-Instr
More results

MM-VQA comprises 2591 colour fundus photographs and 19,648 question-answer pairs.

For TFQ, the fine-tuned model reached an accuracy of 0.919, outperforming Gemini 3 Pro (0.881), Qwen3-VL-30B-A3B-Instruct (0.834), Claude Sonnet 4.5 (0.796), and the pre-trained model (0.696) (all P < 0.001).

More results

On OEQ, it also ranked highest (0.572), outperforming Gemini 3 Pro (0.567, P = 0.044), Claude Sonnet 4.5 (0.395, P < 0.001), Qwen3-VL-30B-A3B-Instruct (0.297, P < 0.001) and the pre-trained model (0.160, P < 0.001).

“
Conclusion

This study provides a valuable visual question answering dataset for myopic maculopathy, supporting the development of disease-specialised vision-language models in ophthalmology.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
0·962 for OCT
MerMED-FM was very accurate at diagnosing diseases using eye scans
AI / Informatics
0.98
Artificial intelligence's eyelid-height measurements matched doctors' manual measurements almost perfectly
AI / Informatics
92.1%
combining eye scans and photos correctly told benign from cancerous lesions apart nearly every time
Cohort Study
28.6%
recurred locally in more than 1 in 4 patients over years of follow-up
Cohort Study
-0.395
excision group's post-op eyelid fullness score was lower - greater improvement
Cohort Study
86.7%
of infants probed after 12 months still had unresolved tear duct blockage
Observational
16%
eyelid tissue in rosacea patients showed less of this key repair-signaling protein inside cell nuclei
Cohort Study
25%
about 1 in 4 treated cases had symptoms return after improving