Multi-reader, multi-model benchmark of large language models for modified Outerbridge cartilage grading from knee MRI reports.
GPT-5.4 matched human radiologists' agreement in grading knee cartilage lesions.
Multi-reader, multi-model benchmark of large language models for modified Outerbridge cartilage grading from knee MRI reports.
To describe a reproducible framework for bulk large language model (large language model)-based extraction of structured cartilage-lesion data from knee MRI reports and to benchmark seven large language model configurations against multiple radiologists using the modified Outerbridge classification.
In this IRB-approved retrospective study, 100 non-contrast knee MRI reports (January 2019 to January 2025) were randomly selected from 66,479 eligible examinations and independently graded by five readers (four fellowship-trained musculoskeletal radiologists with 6-21 years of post-fellowship experience and one fourth-year resident) and seven large language model configurations, comprising six Azure OpenAI deployments (GPT-4.1, GPT-5.1-mini, GPT-5.3, GPT-5.4, GPT-5.4-mini, GPT-5.4-nano) and one locally hosted open-weight model (Qwen2.5-32B-Instruct), across a fixed 20-surface anatomic taxonomy.
GPT-5.4's grading agreement with radiologists matched agreement among radiologists themselves
Four models had CIs entirely below zero, indicating agreement significantly below the human reference: GPT-4.1, GPT-5.4-mini, Qwen2.5-32B, and GPT-5.4-nano (all p ≤ .02).
Three flagship OpenAI deployments achieved cartilage grading agreement indistinguishable from human inter-rater variability, while cost-optimized and open-weight variants performed measurably below.