Post

New research · Ophthalmology
Translational vision science & technology · 2d
AI / informaticsTranslational vision science & technology · 2026

Leakage-Aware Visit-Level Benchmarking Reveals Representational Overlap in Deep Learning for Bacterial Versus Fungal Keratitis Classification.

Wonjong Jeong, Jinho Jeong
Read paper
OphthalmologyAI / informatics

AI struggles to distinguish bacterial from fungal corneal infections due to image overlap

Leakage-Aware Visit-Level Benchmarking Reveals Representational Overlap in Deep Learning for Bacterial Versus Fungal Keratitis Classification.

Wonjong Jeong, Jinho Jeong
Translational vision science & technology · 2026
Purpose

The purpose of this study was to benchmark deep learning for bacterial versus fungal keratitis classification from slit-lamp photographs under a leakage-aware, visit-level framework, and to characterize embedding-space geometry as a potential performance ceiling.

Methods

This retrospective single-center study included white-light slit-lamp photographs from 101 patients with culture-confirmed bacterial or fungal keratitis (258 visits; 658 images), using strict patient-disjoint five-fold cross-validation with the clinical visit as the primary prediction unit.

n = 658 images
90.9%
Results
90.9%
91% of fungal infection images had features overlapping with bacterial ones, confusing AI models.
n = 658 images
More results

At the image level, the highest area under the receiver operating characteristic curve (AUROC) was 0.691.

At the visit level, convolutional neural network (convolutional neural network)-based multiple instance learning models achieved the highest AUROC (0.677), followed by retrieval-based approaches (0.661).

More results

However, multiple instance learning models showed larger generalization gaps, whereas DINOv2 lesion-crop retrieval remained stable (-0.014).

Lesion-centered preprocessing was the dominant performance determinant within the retrieval paradigm.

“
Conclusion · 1 of 3

Under leakage-aware, visit-level evaluation, all paradigms converged to modest discrimination, consistent with representational overlap within this single-center cohort.

Conclusion · 2 of 3

convolutional neural network-based multiple instance learning achieved higher AUROC but greater overfitting; retrieval-based models provided more stable generalization and superior probability calibration.

Conclusion · 3 of 3

Whether this ceiling is task-intrinsic or site-specific requires multicenter validation.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Cross-sectional
44.1%
had too much homocysteine in the blood, a far larger share than in dry AMD or healthy eyes
Cohort Study
92.3%
most laser-treated eyes had a favorable eye-structure outcome
Cohort Study
56%
lost three or more lines on the eye-chart vision test
Cohort Study
50%
older patients were more likely to lose three or more lines of vision
Study
81%
trials where mask air leaks were detected blowing toward the eyes
Cohort Study
13.5%
Central retinal artery peak blood-flow speed was 13.5% lower in diabetic patients.
Study
379 +/- 156 microm
lower thickness in the central retina, checked with an eye scan
Case Report
9 of 10 eyes
eyes had better measured vision after treatment