Post

New research · Cardiology
NPJ digital medicine · 1d
AI / informaticsNPJ digital medicine · 2026

Psychometric characterization of human and artificial intelligence performance on cardiology residency in-service examination items.

Aykan Çelik, Tuncay Kırış, Uğur Kocabaş … Mustafa Karaca
Read paper
CardiologyAI / informatics

A frontier large language model achieved 86.4% accuracy on cardiology residency exam items.

Psychometric characterization of human and artificial intelligence performance on cardiology residency in-service examination items.

Aykan Çelik et al. · NPJ digital medicine · 2026
Background

Large language models (LLMs) are increasingly evaluated using medical examination datasets, yet most studies emphasize overall accuracy rather than the psychometric structure of test items.

Methods

We evaluated five LLMs on 199 text-only cardiology residency in-service examination items previously characterized using resident-derived psychometric metrics.

n = 199 text-only cardiology residency
86.4%
Results
The AI correctly answered this many heart doctor exam questions
n = 199 text-only cardiology residency
More results

Three frontier models were compared with two open-source comparators using a standardized zero-shot, repeated-query protocol and strict-majority scoring.

Across frontier models, performance increased progressively from hard to easy resident-derived item strata.

More results

IRT-based re-analysis confirmed that higher latent item difficulty was independently associated with lower frontier-model accuracy.

These findings show that item-level psychometric analysis helps explain variation in frontier LLM performance beyond overall accuracy alone.

“
Conclusion

However, expert-rated rationale assessment and none-of-the-above perturbation testing revealed that strong examination accuracy did not guarantee high-quality explanatory support or reliable recognition of answer absence, indicating that examination performance and answer-selection robustness represent related but distinct dimensions of model behavior.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Study
89.52%
Nearly 9 in 10 women in cardiology faced bullying, discrimination, or harassment at work
Cross-sectional
18.2%
Roughly 1 in 5 athletes showed a thicker-walled heart geometry pattern on echocardiogram.
Cohort Study
5.8%
Roughly 1 in 17 patients had a cardiovascular death, heart attack, stroke, or unstable angina
Cohort Study
HR 2.02
Highest third of amplified P-wave duration had ~2x the rate of new atrial fibrillation
Study
up to 17.5 mm
the maximum distance the heart treatment area's center moved during breathing
Cohort Study
ranging from -0.13% to 2.70%
percentage of patients whose risk of death changed, often to higher risk
AI / Informatics
F1 = 0.88
this F1 score indicates high accuracy for artificial intelligence models detecting irregular heartbeat
Cohort Study
OR = 2.78
medium triglyceride-glucose index means much higher odds of serious heart problems during hospitalization