Post

New research · AI in Medicine
Journal of the American Medical Informatics Association : JAMIA · 1w
Cohort studyJournal of the American Medical Informatics Association : JAMIA · 2026

Characterization and Validation of EHR Computable Phenotypes for Long COVID Using Patient-Reported Symptoms: Insights from the Nationwide RECOVER Program.

Victor M Castro, Vivian Gainer, Nich Wattanasin … Shawn N Murphy
Read paper
AI in MedicineCohort study

Machine-learning model using electronic health records accurately identified highly symptomatic long COVID

Characterization and Validation of EHR Computable Phenotypes for Long COVID Using Patient-Reported Symptoms: Insights from the Nationwide RECOVER Program.

Victor M Castro … Shawn N Murphy
Journal of the American Medical Informatics Association : JAMIA · 2026
Background

Long COVID (Long COVID) remains poorly understood, and there is a critical need for advanced computational tools to better identify and characterize patients.

Purpose

Long COVID (Long COVID) remains poorly understood, and there is a critical need for advanced computational tools to better identify and characterize patients.

Methods

MAIN OUTCOME AND MEASURES: We assessed model discrimination and calibration in a held-out test set.

n = 1,501 RECOVER-Adult cohort
Results

Model's accuracy score (AUROC) for spotting long COVID; 1.0 is perfect, 0.80 is good

AUROC 0.80 (95% CI 0.74 to 0.85)
null 0
0.74
0.85
CI excludes the null - significant
More results

The study included 1,501 RECOVER-Adult cohort participants with linked EHR data.

376 (25%) met criteria for highly symptomatic Long COVID based on the RECOVER Long COVID Research Index (Long COVID Research Index).

More results

EHR features associated with Long COVID included clinician diagnosis of shortness of breath, malaise and fatigue, and cardiac dysrhythmias; documented treatment with albuterol, gabapentin, or duloxetine; or elevated heart rate.

“
Conclusion · 1 of 2

These findings demonstrate that, using EHR data, a machine-learning model can accurately select patients with sets of self-reported Long COVID symptoms.

Conclusion · 2 of 2

The model could help identify patients within a health system with the highest probability of the condition and facilitate screening, recruitment for clinical trials, and etiologic studies.

Read paper
0 comments

No comments yet. Be the first.

Related papers

LatestFoundational
Randomized Trial
88.2%
Model correctly classified 88.2% of children as achieving remission or not
Cohort Study
75.7%
This is the percentage of people with dementia that the sleep model correctly identified
Guideline
first 5 years
annual eye screening for damage can be delayed for this long if no risk factors
Cohort Study
AUC = 0.856
how well the model separated infants whose own liver survived 2 years from those who needed a transplant; higher means b
AI / Informatics
0.80
the AI's calls on who was at high psychosis risk were mostly correct
AI / Informatics
89%
Artificial intelligence/machine learning models correctly identified 89% of high-grade glioma recurrences and non-recurrences.
Cohort Study
4 percentage points
the new method improved the accuracy of survival prediction for rare cancers
AI / Informatics
88.5%
Top-5 genetic diagnoses were more accurate with Retina4IRD-assisted specialists.