Large language model accuracy in dental radiology: effects of cognitive complexity and content domain.
Low cognitive complexity questions had 6.15 times higher odds of correct answers.
Large language model accuracy in dental radiology: effects of cognitive complexity and content domain.
Large language models (LLMs) are increasingly used to answer medical questions; however, their performance may vary depending on task characteristics.
This study evaluated the performance of multiple versions of two widely used LLM families on oral and maxillofacial radiology (OMFR) questions from the Turkish Dental Specialty Examination (DUS) across three assessment phases and examined the influence of cognitive complexity, model family, evaluation phase, and content domain on response accuracy.
A comparative repeated-evaluation design was used.
A total of 123 text-based OMFR questions from DUS examinations (2012-2021) were submitted to two widely used LLM families (ChatGPT and DeepSeek) across three evaluation phases (May 2025, August 2025, and February 2026).
Agreement between repeated runs was substantial to almost perfect (κ = 0.689-0.912).
Content domain was also associated with accuracy (p = 0.028), whereas no statistically significant associations were observed for model family or evaluation phase.
LLMs demonstrated high accuracy in answering OMFR examination questions; however, performance was more strongly associated with cognitive complexity than with model family or evaluation phase.