Post

New research · Radiology
Scientific reports · 1d
AI / informaticsScientific reports · 2026

Large language model accuracy in dental radiology: effects of cognitive complexity and content domain.

Emre Sözen
Read paper
RadiologyAI / informatics

Low cognitive complexity questions had 6.15 times higher odds of correct answers.

Large language model accuracy in dental radiology: effects of cognitive complexity and content domain.

Emre Sözen · Scientific reports · 2026
Background

Large language models (LLMs) are increasingly used to answer medical questions; however, their performance may vary depending on task characteristics.

Purpose

This study evaluated the performance of multiple versions of two widely used LLM families on oral and maxillofacial radiology (OMFR) questions from the Turkish Dental Specialty Examination (DUS) across three assessment phases and examined the influence of cognitive complexity, model family, evaluation phase, and content domain on response accuracy.

Results
Large language models were much more likely to answer easier questions correctly than harder ones
n = 123 text-based OMFR questions
More results

A comparative repeated-evaluation design was used.

A total of 123 text-based OMFR questions from DUS examinations (2012-2021) were submitted to two widely used LLM families (ChatGPT and DeepSeek) across three evaluation phases (May 2025, August 2025, and February 2026).

More results

Agreement between repeated runs was substantial to almost perfect (κ = 0.689-0.912).

Content domain was also associated with accuracy (p = 0.028), whereas no statistically significant associations were observed for model family or evaluation phase.

“
Conclusion

LLMs demonstrated high accuracy in answering OMFR examination questions; however, performance was more strongly associated with cognitive complexity than with model family or evaluation phase.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Review
10-60 mGy
the range of radiation exposure to the baby from belly or pelvis CT scans
Cross-sectional
56.5%
More than half of computed tomography reports did not mention tumor deposits.
Cohort Study
6.7%
Among hereditary hemorrhagic telangiectasia patients re-imaged, 6.7% had clinically significant liver arteriovenous malformations.
Cohort Study
kappa = 0.819-0.838
adjudicated CT reports showed excellent agreement with the final diagnosis
AI / Informatics
0.87 million
the AI model learned from nearly a million image-report pairs
Guideline
“In those instances where peer reviewed literature is lacking or equivocal, experts may be the primary evidentiary source available to formulate a recommendation.”
Study
5.24 million
the number of brain scans used to train the new artificial intelligence model
AI / Informatics
106 anatomical structures
whole-body reference charts were created for this many different body parts