Large language models accurately identified clinical high risk for psychosis from interview transcripts
Evaluating large language models for assessment of psychosis risk.
Taiyu Zhu … Dominic Oliver
NPJ digital medicine · 2026
Background
Psychosis prevention relies on early detection of individuals at clinical high risk for psychosis (clinical high risk for psychosis).
Methods
We assessed 11 open-weight LLMs on 678 partial PSYCHS interview transcripts from 373 participants (77.7% clinical high risk for psychosis).
n = 373 participants
0.80
Results
0.80
the AI's calls on who was at high psychosis risk were mostly correct
n = 373 participants
More results
Models inferred clinical high risk for psychosis status and estimated severity and frequency across 15 symptom domains, benchmarked against researcher-rated scores.
LLM-generated symptom scores showed good correlations with researcher-rated scores (ICC sev = 0.74, ICC freq = 0.75).
More results
Generated summaries were largely faithful to source transcripts, with low rates of clinically relevant confabulation (3%).
While accuracy scaled with model size, smaller models achieved competitive performance with substantially lower computational cost.
“
Conclusion
These findings demonstrate that open-weight LLMs have the potential to assess psychosis risk from psychometric interview transcripts, supporting scalable, human-in-the-loop approaches to early detection.