A machine learning-based risk prediction model for Hospitalized patients with deep vein thrombosis.
Random Forest model predicted deep vein thrombosis risk with strong discrimination.
A machine learning-based risk prediction model for Hospitalized patients with deep vein thrombosis.
Deep vein thrombosis (deep vein thrombosis) is a common thrombotic condition with substantial morbidity when not identified early.
To develop and internally validate a machine learning model using routinely available clinical and laboratory indicators for early risk prediction of deep vein thrombosis, and to identify the most influential predictors using model explainability techniques.
We retrospectively analyzed clinical data from 231 patients evaluated at the Fifth Affiliated Hospital of Southern Medical University between January 2017 and June 2024.
Least Absolute Shrinkage and Selection Operator selected seven predictors: hemoglobin, platelet count, leukocyte count, fibrinogen, prothrombin time, D-dimer (d-dimer), and glucose.
In the Random Forest model, D-dimer had the highest feature-importance contribution; SHAP analysis confirmed d-dimer as the dominant risk driver and characterized the directions and relative effects of other features.
We developed an internally validated machine learning model for early deep vein thrombosis risk prediction using seven routine clinical variables; Random Forest achieved the best performance and identified D-dimer as the most influential predictor.
This model may support earlier identification and intervention for patients at risk of deep vein thrombosis, pending external validation and prospective evaluation.