Post

New research · Psychiatry
Canadian journal of psychiatry. Revue canadienne de psychiatrie · 1d
AI / informaticsCanadian journal of psychiatry. Revue canadienne de psychiatrie · 2026

Zero-Shot Large Language Models for Preliminary Prediction of PTSD Symptoms From Clinical Interview Transcripts: Grands modèles de langage sans exemple pour la prédiction préliminaire des symptômes de TSPT à partir de transcriptions d'entrevues cliniques.

Bazen Gashaw Teferra, Christian Kevin Sidharta, Wei-Ni Hsiang … Venkat Bhat
Read paper
PsychiatryAI / informatics

Claude 4 achieved 0.705 accuracy predicting binary posttraumatic stress disorder symptoms from interviews.

Zero-Shot Large Language Models for Preliminary Prediction of PTSD Symptoms From Clinical Interview Transcripts: Grands modèles de langage sans exemple pour la prédiction préliminaire des symptômes de TSPT à partir de transcriptions d'entrevues cliniques.

Bazen Gashaw Teferra et al. · Canadian journal of psychiatry. Revue canadienne de psychiatrie · 2026
Background

Posttraumatic stress disorder (PTSD) is common yet frequently underdiagnosed, in part due to barriers to systematic screening and the reliance on self-report instruments.

Methods

Using the Distress Analysis Interview Corpus-Wizard of Oz (DAIC-WoZ), we analyzed 100 semi-structured clinical interview transcripts paired with item-level PTSD Checklist-Civilian Version (PCL-C) scores.

n = 100 semi-structured clinical
Results

this is how accurately Claude 4 predicted if someone had PTSD symptoms

mean accuracy 0.70 (95% CI 0.68 to 0.73)
null = 00.680.73
CI excludes the null - significant
More results

For Likert prediction, DeepSeek 3.1 performed best (accuracy = 0.438; 95% CI, 0.401-0.475), only modestly above the majority-class baseline (0.399; 95% CI, 0.355-0.443).

Across models, predicted item-level symptom patterns showed a meaningful alignment with observed PCL-C responses despite reduced accuracy in fine-grained severity estimation.

“
Conclusion · 1 of 3

Zero-shot LLMs' performance was insufficient for clinical application in predicting PTSD symptoms from semi-structured interview transcripts.

Conclusion · 2 of 3

While models showed some ability to capture overall symptom patterns, performance varied across domains and remained limited for fine-grained severity estimation.

Conclusion · 3 of 3

Given these constraints and the non-trauma-specific nature of the dataset, findings should be interpreted as preliminary, with only modest differences observed between models.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
AI / Informatics
0.89-0.93
the high level of agreement between the AI rater and psychiatrists on consultation sections
Cohort Study
95.5%
patients identified as low risk by the model did not experience psychiatric worsening
AI / Informatics
0.72
This accuracy placed top LLMs at the 64th percentile of clinician performance.
Guideline
“To enhance local relevance, countries should adapt recommendations to national service and policy contexts.”
Cross-sectional
approximately 50%
participants reported less severe depression symptoms
Guideline
“Clinicians should not use parenteral pharmacological prophylaxis in adults hospitalised with psychiatric illness at low risk of VTE; and should consider using parenteral pharmacological prophylaxis for high-risk adults with no contraindications.”
Case Report
4 weeks
delusions, paranoia, and thoughts of self-harm went away in this time
Systematic Review
96%-99%
biomarker- and cellular-based Artificial Intelligence models correctly predict long-term treatment success