Leakage-Aware Visit-Level Benchmarking Reveals Representational Overlap in Deep Learning for Bacterial Versus Fungal Keratitis Classification.
AI struggles to distinguish bacterial from fungal corneal infections due to image overlap
Leakage-Aware Visit-Level Benchmarking Reveals Representational Overlap in Deep Learning for Bacterial Versus Fungal Keratitis Classification.
The purpose of this study was to benchmark deep learning for bacterial versus fungal keratitis classification from slit-lamp photographs under a leakage-aware, visit-level framework, and to characterize embedding-space geometry as a potential performance ceiling.
This retrospective single-center study included white-light slit-lamp photographs from 101 patients with culture-confirmed bacterial or fungal keratitis (258 visits; 658 images), using strict patient-disjoint five-fold cross-validation with the clinical visit as the primary prediction unit.
At the image level, the highest area under the receiver operating characteristic curve (AUROC) was 0.691.
At the visit level, convolutional neural network (convolutional neural network)-based multiple instance learning models achieved the highest AUROC (0.677), followed by retrieval-based approaches (0.661).
However, multiple instance learning models showed larger generalization gaps, whereas DINOv2 lesion-crop retrieval remained stable (-0.014).
Lesion-centered preprocessing was the dominant performance determinant within the retrieval paradigm.
Under leakage-aware, visit-level evaluation, all paradigms converged to modest discrimination, consistent with representational overlap within this single-center cohort.
convolutional neural network-based multiple instance learning achieved higher AUROC but greater overfitting; retrieval-based models provided more stable generalization and superior probability calibration.
Whether this ceiling is task-intrinsic or site-specific requires multicenter validation.