Reporting Quality of Large Language Model Studies: A Cross-Sectional Audit of High-Ranking Radiology and Medical Imaging Journals.
Radiology large language models studies show only moderate reporting-guideline adherence, averaging 51.2%.
Reporting Quality of Large Language Model Studies: A Cross-Sectional Audit of High-Ranking Radiology and Medical Imaging Journals.
To evaluate adherence to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) in radiology and medical imaging studies involving large language models (large language models).
We conducted a cross-sectional audit of original large language models research studies published between January 1 and December 26, 2025, in Q1 journals within the Web of Science "Radiology, Nuclear Medicine, and Medical Imaging" category.
Of 201 eligible studies identified, 102 were finally analyzed after applying the subsampling strategy.
Adherence was highest for input data type (100%), test-data independence (80.2%), and adaptation strategy (78.1%), and lowest for prompt execution setup (29.4%) and stochasticity management (33.1%).
The least frequently reported items were training-data cutoff date (9.8%) and rationale for prompt wording (15.6%).
Adherence varied significantly across journals ( P = 0.011), with KJR showing the highest mean adherence (72.8% ± 2.7%).
Reporting transparency in radiology and medical imaging large language models studies published in 2025 was inconsistent across reporting items and journals, with substantial deficiencies in some reproducibility-critical elements.
Broader adoption of reporting standards is essential to improve the reproducibility and interpretability of future accuracy evaluations.