Post

New research · AI in Medicine
Journal of psychopathology and clinical science · 1w
AI / informaticsJournal of psychopathology and clinical science · 2026

Explaining GPTs' schema of depression: A machine behavior analysis.

Adithya V Ganesan, Vasudha Varadarajan, Yash Kumar Lal … Lucie Flek
Read paper
AI in MedicineAI / informatics

GPT-4's depression assessments strongly agreed with standard instruments and expert judgments

Explaining GPTs' schema of depression: A machine behavior analysis.

Adithya V Ganesan … Lucie Flek
Journal of psychopathology and clinical science · 2026
Methods

First, we evaluated whether these models can reliably detect depressive constructs, finding that GPT-4 (a) demonstrated strong convergent validity with standard instruments and expert judgments ( r = .70-.81).

r = .70-.81
Results
r = .70-.81
GPT-4's depression ratings closely matched standard questionnaires and expert clinicians' judgments
More results

The use of large language models (large language models) such as ChatGPT (GPT-4/GPT-5) for mental health is growing rapidly, including their use to assess and help people with mood disorders, such as depression.

More results

In this work, we used contemporary measurement theory to decode how GPT-4 and GPT-5 interrelate depressive symptoms, providing an explanation of how large language models behave in clinical applications.

“
Conclusion

(PsycInfo Database Record (c) 2026 APA, all rights reserved).

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Cohort Study
AUC = 0.856
how well the model separated infants whose own liver survived 2 years from those who needed a transplant; higher means b
Guideline
“This approach enables improved data consistency, facilitates national health information exchange, and lays a foundation for advanced applications, such as clinical decision support and artificial intelligence-driven analytics.”
Randomized Trial
88.2%
Model correctly classified 88.2% of children as achieving remission or not
AI / Informatics
0.80
the AI's calls on who was at high psychosis risk were mostly correct
Cohort Study
0.80
Model's accuracy score (AUROC) for spotting long COVID; 1.0 is perfect, 0.80 is good
AI / Informatics
88.5%
Top-5 genetic diagnoses were more accurate with Retina4IRD-assisted specialists.
AI / Informatics
27,624 words
total Simplified Chinese words with new AI familiarity estimates now available
AI / Informatics
89%
Artificial intelligence/machine learning models correctly identified 89% of high-grade glioma recurrences and non-recurrences.