Post

New research · AI in Medicine
NPJ digital medicine · 1w
AI / informaticsNPJ digital medicine · 2026

Evaluating large language models for assessment of psychosis risk.

Taiyu Zhu, Alexander Tashevski, Maxime Taquet … Dominic Oliver
Read paper
AI in MedicineAI / informatics

Large language models accurately identified clinical high risk for psychosis from interview transcripts

Evaluating large language models for assessment of psychosis risk.

Taiyu Zhu … Dominic Oliver
NPJ digital medicine · 2026
Background

Psychosis prevention relies on early detection of individuals at clinical high risk for psychosis (clinical high risk for psychosis).

Methods

We assessed 11 open-weight LLMs on 678 partial PSYCHS interview transcripts from 373 participants (77.7% clinical high risk for psychosis).

n = 373 participants
0.80
Results
0.80
the AI's calls on who was at high psychosis risk were mostly correct
n = 373 participants
More results

Models inferred clinical high risk for psychosis status and estimated severity and frequency across 15 symptom domains, benchmarked against researcher-rated scores.

LLM-generated symptom scores showed good correlations with researcher-rated scores (ICC sev = 0.74, ICC freq = 0.75).

More results

Generated summaries were largely faithful to source transcripts, with low rates of clinically relevant confabulation (3%).

While accuracy scaled with model size, smaller models achieved competitive performance with substantially lower computational cost.

“
Conclusion

These findings demonstrate that open-weight LLMs have the potential to assess psychosis risk from psychometric interview transcripts, supporting scalable, human-in-the-loop approaches to early detection.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
0.32
AUC was 0.32 higher, meaning better separation of responders from non-responders
Review
41.9%
nearly half of people with a failing heart also experience depression
Cohort Study
4 percentage points
the new method improved the accuracy of survival prediction for rare cancers
AI / Informatics
0.72
This accuracy placed top LLMs at the 64th percentile of clinician performance.
AI / Informatics
0.64-0.69
The modest ability of AI-ECG alone to predict irregular heartbeat risk
AI / Informatics
0.987
model's accuracy in identifying Cadmium contamination levels
AI / Informatics
5,188
the total number of expert-reviewed questions used to test AI models
Cross-sectional
approximately 50%
participants reported less severe depression symptoms