Large language models show low accuracy for case-based clinical scenarios.
ChatGPT and Other Large Language Models in Inflammatory Arthritis: A Systematic Review Across Clinical Tasks.
Eighteen studies covered rheumatoid arthritis (n=3), ankylosing spondylitis/axial spondyloarthritis (n=7), psoriatic arthritis (n=2), gout (n=1), juvenile idiopathic arthritis (n=1), and multiple diseases (n=4).
Most diseases and tasks were represented by only one to a few studies, and the evidence base remains earlystage and uneven across conditions.
Guideline concordance ranged from 48% to 96%.
When compared with real clinical data, agreement was poor (Cohen and Fleiss κ ≈ 0).
LLMs may support patient education, factual medication queries, and structured guideline questions when used under clinician review, but should not be used for case-based reasoning, treatment selection, or autonomous clinical decisions.
None of the 18 included studies evaluated retrieval-augmented or agent-based systems, and none prospectively validated LLMs in clinical workflows.
Safe integration in rheumatology will require purpose-built, knowledge-grounded systems and prospective evaluation before routine clinical use.