PsychiatryBench is a new benchmark with 5,188 expert-annotated items for LLM evaluation.
PsychiatryBench: a multi-task benchmark for LLMs in psychiatry.
Aya E Fouda et al. · NPJ digital medicine · 2026
Background
Large language models (LLMs) offer significant potential in enhancing psychiatric practice, from improving diagnostic accuracy to streamlining clinical documentation and therapeutic support.
Results
the total number of expert-reviewed questions used to test AI models
n = 5,188 expert-annotated items
More results
We evaluate a diverse set of frontier LLMs (including Google Gemini, DeepSeek, Sonnet 4.5, and GPT 5) alongside leading open-source medical models such as MedGemma using both conventional metrics and an "LLM-as-judge" similarity scoring framework.
“
Conclusion
PsychiatryBench offers a modular, extensible platform for benchmarking and improving LLM performance in mental health applications.