Post

New research · Psychiatry
NPJ digital medicine · 2d
AI / informaticsNPJ digital medicine · 2026

PsychiatryBench: a multi-task benchmark for LLMs in psychiatry.

Aya E Fouda, Abdelrahman A Hassan, Radwa J Hanafy, Mohammed E Fouda
Read paper
PsychiatryAI / informatics

PsychiatryBench is a new benchmark with 5,188 expert-annotated items for LLM evaluation.

PsychiatryBench: a multi-task benchmark for LLMs in psychiatry.

Aya E Fouda et al. · NPJ digital medicine · 2026
Background

Large language models (LLMs) offer significant potential in enhancing psychiatric practice, from improving diagnostic accuracy to streamlining clinical documentation and therapeutic support.

Results
the total number of expert-reviewed questions used to test AI models
n = 5,188 expert-annotated items
More results

We evaluate a diverse set of frontier LLMs (including Google Gemini, DeepSeek, Sonnet 4.5, and GPT 5) alongside leading open-source medical models such as MedGemma using both conventional metrics and an "LLM-as-judge" similarity scoring framework.

“
Conclusion

PsychiatryBench offers a modular, extensible platform for benchmarking and improving LLM performance in mental health applications.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Review
41.9%
nearly half of people with a failing heart also experience depression
AI / Informatics
0.72
This accuracy placed top LLMs at the 64th percentile of clinician performance.
Cohort Study
95.5%
patients identified as low risk by the model did not experience psychiatric worsening
Cross-sectional
approximately 50%
participants reported less severe depression symptoms