Post

New research · Ophthalmology
Veterinary ophthalmology · 23h
Animal / preclinicalVeterinary ophthalmology · 2026

Challenges and Misinterpretations of Cohen's Kappa in Agreement Studies in Ophthalmology.

Malwina Ewa Kowalska, Niklas Holz, Simon Anton Pot … Sonja Hartnack
Read paper
OphthalmologyAnimal / preclinical

Balancing data distribution substantially increased Cohen's kappa agreement for left eyes.

Challenges and Misinterpretations of Cohen's Kappa in Agreement Studies in Ophthalmology.

Malwina Ewa Kowalska et al. · Veterinary ophthalmology · 2026
Background

Cohen's kappa is widely used to measure rater agreement on binary outcomes and performs well when outcome prevalence is relatively balanced.

Purpose

This paper aims to explain the challenges encountered with Cohen's kappa, and to provide guidance to veterinary ophthalmologists on the usage and interpretation of kappa values.

Methods

Following the 2022 ECVO-HED (European College of Veterinary Ophthalmology-Hereditary Eye Diseases) gonioscopy grading scheme and its general recommendations, eyes were classified into two categories-"breeding-YES" (not affected, mildly or moderately affected) or "breeding-NO" (severely affected).

n = 60 eyes
0.77
Results
This shows higher agreement between examiners on left eye condition after balancing the data
n = 60 eyes
More results

While the two examiners classified 52/60 left and 51/58 right eyes into the "breeding-YES" category, they disagreed on 7/60 left and 1/58 right eyes, with resulting kappa values of 0.18 and 0.91 for the left and right eyes, respectively.

“
Conclusion · 1 of 2

Readers should be cautious when comparing kappa values across studies with different underlying prevalence and bias.

Conclusion · 2 of 2

Under non-ideal data distribution, additional statistical indices like prevalence- and bias index, prevalence-and-bias-adjusted kappa (PABAK), and maxKappa, do not replace Cohen's kappa, but assist in interpreting it.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
difference, 0
AI and human experts had the same median quality scores for eye care questions
Case-control
0 of 28
distinct structural changes were absent in all 28 healthy control eyes
Cohort Study
RR, 0.91
Black patients were less likely to start medication for their eye condition
AI / Informatics
100% diagnostic sensitivity
ROFI correctly found every case of eye disease
AI / Informatics
10.50-point increase
students trained with digital patients had higher history-taking assessment scores
AI / Informatics
0.877
The o1 large language model correctly answered 87.7% of ophthalmology questions.
Cohort Study
HR 4.85
chronic pain conditions mean a nearly five times higher risk of developing dry eye disease
Cohort Study
12 per 10,000
a higher rate of eye infections after dexamethasone implant injections