ChatGPT and Other Large Language Models in Inflammatory Arthritis: A Systematic Review Across Clinical Tasks.
Large language models show low accuracy for case-based clinical scenarios.
ChatGPT and Other Large Language Models in Inflammatory Arthritis: A Systematic Review Across Clinical Tasks.
Large language models (large language models) are increasingly evaluated for rheumatology tasks, but their performance in inflammatory arthritis remains unclear.
Large language models (large language models) are increasingly evaluated for rheumatology tasks, but their performance in inflammatory arthritis remains unclear.
We conducted a systematic review (PROSPERO: CRD420261359100), searching PubMed, Scopus, and PubMed Central (January 2022 to April 2026) for studies evaluating large language models performance on clinical tasks in inflammatory arthritis.
Eighteen studies covered rheumatoid arthritis (n=3), ankylosing spondylitis/axial spondyloarthritis (n=7), psoriatic arthritis (n=2), gout (n=1), juvenile idiopathic arthritis (n=1), and multiple diseases (n=4).
Most diseases and tasks were represented by only one to a few studies, and the evidence base remains earlystage and uneven across conditions.
Guideline concordance ranged from 48% to 96%.
When compared with real clinical data, agreement was poor (Cohen and Fleiss κ ≈ 0).
Large language models may support patient education, factual medication queries, and structured guideline questions when used under clinician review, but should not be used for case-based reasoning, treatment selection, or autonomous clinical decisions.
None of the 18 included studies evaluated retrieval-augmented or agent-based systems, and none prospectively validated large language models in clinical workflows.
Safe integration in rheumatology will require purpose-built, knowledge-grounded systems and prospective evaluation before routine clinical use.