Post

New research · Psychiatry
NPJ digital medicine · 2d
AI / informaticsNPJ digital medicine · 2026

Benchmarking large language models against practicing clinicians on psychopathological assessment.

Esra Lenz, Joonas Naamanka, Wolfgang Trabert … Emanuel Schwarz
Read paper
PsychiatryAI / informatics

Top LLMs achieved 0.72 accuracy in psychiatric symptom assessment.

Benchmarking large language models against practicing clinicians on psychopathological assessment.

Esra Lenz et al. · NPJ digital medicine · 2026
Background

Psychiatry's reliance on language makes LLMs a natural tool for psychopathological assessment, yet structured, item-level assessments from psychiatric clinical interviews remain under-researched.

Methods

These proof-of-concept findings require validation in real patient interviews, larger samples, and prospective studies integrating multimodal input.

0.72
Results
This accuracy placed top LLMs at the 64th percentile of clinician performance.
More results

In this proof-of-concept study, 10 LLMs assessed transcripts of three simulated psychiatric interviews across all 100 items of the Association for Methodology and Documentation in Psychiatry (AMDP) system, benchmarked against 108 early-career clinicians rating full audiovisual recordings, using an expert consensus panel as reference.

More results

GPT-5.1, selected for a marginal advantage, showed per-scenario accuracies of 0.81 (depression), 0.76 (mania), and 0.60 (schizophrenia) versus clinician means of 0.79, 0.68, and 0.58.

“
Conclusion

These proof-of-concept findings require validation in real patient interviews, larger samples, and prospective studies integrating multimodal input.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Review
41.9%
nearly half of people with a failing heart also experience depression
AI / Informatics
5,188
the total number of expert-reviewed questions used to test AI models
Cohort Study
95.5%
patients identified as low risk by the model did not experience psychiatric worsening
Cross-sectional
approximately 50%
participants reported less severe depression symptoms