Evaluating the diagnostic performance of multimodal large language models in thyroid cytology.
ChatGPT correctly identified specific thyroid fine-needle aspiration diagnoses only 39.2% of the time
Evaluating the diagnostic performance of multimodal large language models in thyroid cytology.
Thyroid fine-needle aspiration (fine-needle aspiration) cytology is a widely used, cost-effective method for evaluating thyroid nodules.
In this study, we evaluated the diagnostic accuracy and consistency of ChatGPT and Gemini in interpreting 100 thyroid fine-needle aspiration cytology images, comprising 50 cases of papillary thyroid carcinoma (papillary thyroid carcinoma) and 50 non-papillary thyroid carcinoma lesions, including follicular nodular disease, follicular neoplasm, medullary thyroid carcinoma, and anaplastic thyroid carcinoma.
ChatGPT matched the exact cytology diagnosis in just over 1 in 3 images
However, interpretation of thyroid fine-needle aspiration smears remains challenging, particularly in settings with limited access to experienced cytopathologists.
Nevertheless, data regarding their application in cytopathology, including thyroid cytology, remain limited.
Representative images were captured at 40× magnification and reviewed by expert pathologists to confirm the reference diagnoses.
ChatGPT rendered a higher number of suspicious for malignancy and atypia of undetermined significance diagnoses than Gemini.
In conclusion, although ChatGPT and Gemini demonstrated the potential of large language models in cytologic image interpretation and drafting descriptive reports, their diagnostic performance in thyroid fine-needle aspiration cytology was limited, underscoring the current constraints of large language models in this domain.