Post

New research · Rheumatology
The Journal of rheumatology · 21h
AI / informaticsThe Journal of rheumatology · 2026

ChatGPT and Other Large Language Models in Inflammatory Arthritis: A Systematic Review Across Clinical Tasks.

Yosef Adiniaev, Mahmud Omar, Tohar M Timor … Eyal Klang
Read paper
RheumatologyAI / informatics

Large language models show low accuracy for case-based clinical scenarios.

ChatGPT and Other Large Language Models in Inflammatory Arthritis: A Systematic Review Across Clinical Tasks.

Yosef Adiniaev … Eyal Klang
The Journal of rheumatology · 2026
Background

Large language models (large language models) are increasingly evaluated for rheumatology tasks, but their performance in inflammatory arthritis remains unclear.

Purpose

Large language models (large language models) are increasingly evaluated for rheumatology tasks, but their performance in inflammatory arthritis remains unclear.

Methods

We conducted a systematic review (PROSPERO: CRD420261359100), searching PubMed, Scopus, and PubMed Central (January 2022 to April 2026) for studies evaluating large language models performance on clinical tasks in inflammatory arthritis.

n = 113 records
4.24/6
Results
4.24/6
Large language models scored lower on accuracy when answering questions about patient cases
n = 113 records
More results

Eighteen studies covered rheumatoid arthritis (n=3), ankylosing spondylitis/axial spondyloarthritis (n=7), psoriatic arthritis (n=2), gout (n=1), juvenile idiopathic arthritis (n=1), and multiple diseases (n=4).

Most diseases and tasks were represented by only one to a few studies, and the evidence base remains earlystage and uneven across conditions.

More results

Guideline concordance ranged from 48% to 96%.

When compared with real clinical data, agreement was poor (Cohen and Fleiss κ ≈ 0).

“
Conclusion · 1 of 3

Large language models may support patient education, factual medication queries, and structured guideline questions when used under clinician review, but should not be used for case-based reasoning, treatment selection, or autonomous clinical decisions.

Conclusion · 2 of 3

None of the 18 included studies evaluated retrieval-augmented or agent-based systems, and none prospectively validated large language models in clinical workflows.

Conclusion · 3 of 3

Safe integration in rheumatology will require purpose-built, knowledge-grounded systems and prospective evaluation before routine clinical use.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Case Report
creatinine at 6.28 mg/dL
high level of a waste product, indicating severe kidney damage
Guideline
≥80% acceptance
most experts accepted the new treatment recommendations
Study
80% (4/5)
Four out of five patients with refractory systemic lupus erythematosus achieved remission.
Review
73%
patient-reported symptoms were used more often to define ongoing symptoms
Cohort Study
AUC of 0.804
how well the three-gene model identifies Dampness ZHENG
Cohort Study
OR 0.70
Odds of valve progression or new valve disease were similar with prophylaxis.
Study
AUC = 0.847
Good accuracy separating lupus patients from healthy people (higher AUC = better)
AI / Informatics
accuracy 0.731
the computer model correctly identified juvenile arthritis activity levels nearly three-quarters of the time