Post

New research · Cardiology
NPJ digital medicine · 1d
AI / informaticsNPJ digital medicine · 2026

Psychometric characterization of human and artificial intelligence performance on cardiology residency in-service examination items.

Aykan Çelik, Tuncay Kırış, Uğur Kocabaş … Mustafa Karaca
Read paper
CardiologyAI / informatics

A frontier large language model achieved 86.4% accuracy on cardiology residency exam items.

Psychometric characterization of human and artificial intelligence performance on cardiology residency in-service examination items.

Aykan Çelik et al. · NPJ digital medicine · 2026
Background

Large language models (LLMs) are increasingly evaluated using medical examination datasets, yet most studies emphasize overall accuracy rather than the psychometric structure of test items.

Methods

We evaluated five LLMs on 199 text-only cardiology residency in-service examination items previously characterized using resident-derived psychometric metrics.

n = 199 text-only cardiology residency
86.4%
Results
The AI correctly answered this many heart doctor exam questions
n = 199 text-only cardiology residency
More results

Three frontier models were compared with two open-source comparators using a standardized zero-shot, repeated-query protocol and strict-majority scoring.

Across frontier models, performance increased progressively from hard to easy resident-derived item strata.

More results

IRT-based re-analysis confirmed that higher latent item difficulty was independently associated with lower frontier-model accuracy.

These findings show that item-level psychometric analysis helps explain variation in frontier LLM performance beyond overall accuracy alone.

“
Conclusion

However, expert-rated rationale assessment and none-of-the-above perturbation testing revealed that strong examination accuracy did not guarantee high-quality explanatory support or reliable recognition of answer absence, indicating that examination performance and answer-selection robustness represent related but distinct dimensions of model behavior.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Cross-sectional
13.9%
new guidelines make more US adults eligible for cholesterol-lowering medication
Cohort Study
95.4%
most patients had the minimally invasive heart repair procedure work as intended
Cohort Study
5.8%
Roughly 1 in 17 patients had a cardiovascular death, heart attack, stroke, or unstable angina
Cohort Study
5.2
higher annual heart attack or stroke rate for borderline risk people with calcium
Study
89.52%
Nearly 9 in 10 women in cardiology faced bullying, discrimination, or harassment at work
Cohort Study
OR = 2.78
medium triglyceride-glucose index means much higher odds of serious heart problems during hospitalization
Cohort Study
3%
lower incidence of low blood sugar with the simplified protocol
Case Report
12 mm
the brain's center was pushed sideways due to severe swelling