Large language model accuracy in dental radiology: effects of cognitive complexity and content domain.
Low cognitive complexity questions had 6.15 times higher odds of correct answers.
Large language model accuracy in dental radiology: effects of cognitive complexity and content domain.
Large language models (large language models) are increasingly used to answer medical questions; however, their performance may vary depending on task characteristics.
This study evaluated the performance of multiple versions of two widely used large language models families on oral and maxillofacial radiology (oral and maxillofacial radiology) questions from the Turkish Dental Specialty Examination (DUS) across three assessment phases and examined the influence of cognitive complexity, model family, evaluation phase, and content domain on response accuracy.
A comparative repeated-evaluation design was used.
A total of 123 text-based oral and maxillofacial radiology questions from DUS examinations (2012-2021) were submitted to two widely used large language models families (ChatGPT and DeepSeek) across three evaluation phases (May 2025, August 2025, and February 2026).
Agreement between repeated runs was substantial to almost perfect (κ = 0.689-0.912).
Content domain was also associated with accuracy (p = 0.028), whereas no statistically significant associations were observed for model family or evaluation phase.
Large language models demonstrated high accuracy in answering oral and maxillofacial radiology examination questions; however, performance was more strongly associated with cognitive complexity than with model family or evaluation phase.