Post

New research · Ophthalmology
Translational vision science & technology · 2d
AI / informaticsTranslational vision science & technology · 2026

Leakage-Aware Visit-Level Benchmarking Reveals Representational Overlap in Deep Learning for Bacterial Versus Fungal Keratitis Classification.

Wonjong Jeong, Jinho Jeong
Read paper
OphthalmologyAI / informatics

AI struggles to distinguish bacterial from fungal corneal infections due to image overlap

Leakage-Aware Visit-Level Benchmarking Reveals Representational Overlap in Deep Learning for Bacterial Versus Fungal Keratitis Classification.

Wonjong Jeong, Jinho Jeong
Translational vision science & technology · 2026
Purpose

The purpose of this study was to benchmark deep learning for bacterial versus fungal keratitis classification from slit-lamp photographs under a leakage-aware, visit-level framework, and to characterize embedding-space geometry as a potential performance ceiling.

Methods

This retrospective single-center study included white-light slit-lamp photographs from 101 patients with culture-confirmed bacterial or fungal keratitis (258 visits; 658 images), using strict patient-disjoint five-fold cross-validation with the clinical visit as the primary prediction unit.

n = 658 images
90.9%
Results
90.9%
91% of fungal infection images had features overlapping with bacterial ones, confusing AI models.
n = 658 images
More results

At the image level, the highest area under the receiver operating characteristic curve (AUROC) was 0.691.

At the visit level, convolutional neural network (convolutional neural network)-based multiple instance learning models achieved the highest AUROC (0.677), followed by retrieval-based approaches (0.661).

More results

However, multiple instance learning models showed larger generalization gaps, whereas DINOv2 lesion-crop retrieval remained stable (-0.014).

Lesion-centered preprocessing was the dominant performance determinant within the retrieval paradigm.

“
Conclusion · 1 of 3

Under leakage-aware, visit-level evaluation, all paradigms converged to modest discrimination, consistent with representational overlap within this single-center cohort.

Conclusion · 2 of 3

convolutional neural network-based multiple instance learning achieved higher AUROC but greater overfitting; retrieval-based models provided more stable generalization and superior probability calibration.

Conclusion · 3 of 3

Whether this ceiling is task-intrinsic or site-specific requires multicenter validation.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
0·962 for OCT
MerMED-FM was very accurate at diagnosing diseases using eye scans
AI / Informatics
0.98
Artificial intelligence's eyelid-height measurements matched doctors' manual measurements almost perfectly
AI / Informatics
92.1%
combining eye scans and photos correctly told benign from cancerous lesions apart nearly every time
Cohort Study
28.6%
recurred locally in more than 1 in 4 patients over years of follow-up
Cohort Study
-0.395
excision group's post-op eyelid fullness score was lower - greater improvement
Cohort Study
86.7%
of infants probed after 12 months still had unresolved tear duct blockage
Observational
16%
eyelid tissue in rosacea patients showed less of this key repair-signaling protein inside cell nuclei
Cohort Study
25%
about 1 in 4 treated cases had symptoms return after improving