Heart Disease Prediction using Cost-sensitive Random Forest and SHAP

Arya Chova Pratama, Eko Arip Winanto, Beny Beny

Abstract


Machine learning–based studies on heart disease prediction often report high predictive accuracy; however, most rely on relatively small and nearly balanced datasets, such as the Cleveland UCI dataset. Such datasets do not adequately reflect real-world population distributions, where class imbalance is substantial and the consequences of misclassification are asymmetric. Missing a positive case may forfeit the opportunity for early intervention, whereas a false positive can generally be resolved through subsequent clinical assessment. This study proposes a methodological evaluation framework for classifying self-reported cardiovascular risk profiles by integrating Cost-Sensitive Random Forest (CSRF) with SHapley Additive exPlanations (SHAP). The analysis uses the 2022 Behavioral Risk Factor Surveillance System (BRFSS) no_nans dataset released by the Centers for Disease Control and Prevention (CDC) and curated by Pytlak, comprising 246,022 respondents, 39 predictor variables, and a class imbalance ratio of 1:17.31. Eight model configurations were evaluated under an identical hyperparameter tuning budget, including a baseline Random Forest, cost-sensitive variants, resampling approaches, and gradient boosting models as robustness benchmarks. Model performance was assessed using cross-validation, bootstrap confidence intervals, leave-one-state-out validation, cost matrix–based threshold optimization, and calibration and fairness audits. On a test set of 49,205 respondents, the CSRF model with a 1:3 cost ratio reduced false negatives from 2,117 to 1,360, achieving a recall of 49.4% (95% CI: 47.5–51.2%) and an F1-score of 0.483. This F1-score was the highest among the Random Forest and resampling approaches, while Cost-Sensitive XGBoost and Cost-Sensitive LightGBM achieved slightly higher F1-scores, providing additional evidence of model robustness. SHAP analysis identified HadAngina and ChestScan as the most influential predictors. Because these variables are clinically closely related to the target outcome, the model's performance should be interpreted primarily as capturing cross-sectional associations derived from self-reported data rather than as a purely prospective clinical prediction model.

Keywords


cost-sensitive random forest; decision curve analysis; explainable AI; false negative; SHAP

Full Text:

PDF

References


Global Burden of Cardiovascular Diseases and Risks 2023 Collaborators, "Global, Regional, and National Burden of Cardiovascular Diseases and Risk Factors in 204 Countries and Territories, 1990–2023," J. Am. Coll. Cardiol., Vol. 86, No. 22, pp. 2167–2243, 2025, DOI: 10.1016/j.jacc.2025.08.015.

K. Pytlak, "Indicators of Heart Disease (2022 Update)," Kaggle, 2023. [Online]. Available: https://www.kaggle.com/datasets/kamilpytlak/personal-key-indicators-of-heart-disease.

L. Breiman, "Random Forests," Mach. Learn., Vol. 45, No. 1, pp. 5–32, 2001, DOI: 10.1023/A:1010933404324.

G. Lemaître, F. Nogueira, and C. K. Aridas, "Imbalanced-Learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning," J. Mach. Learn. Res., Vol. 18, No. 17, pp. 1–5, 2017.

I. Araf, A. Idri, and I. Chairi, "Cost-Sensitive Learning for Imbalanced Medical Data: A Review," Artif. Intell. Rev., Vol. 57, No. 4, Art. 80, 2024, DOI: 10.1007/s10462-023-10652-8.

S. M. Lundberg, G. Erion, H. Chen, A. DeGrave, J. M. Prutkin, B. Nair, R. Katz, J. Himmelfarb, N. Bansal, and S.-I. Lee, "From Local Explanations to Global Understanding with Explainable AI for Trees," Nat. Mach. Intell., Vol. 2, No. 1, pp. 56–67, 2020, DOI: 10.1038/s42256-019-0138-9.

N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: Synthetic Minority Over-Sampling Technique," J. Artif. Intell. Res., Vol. 16, pp. 321–357, 2002, DOI: 10.1613/jair.953.

M. Zhu, J. Xia, X. Jin, M. Yan, G. Cai, J. Yan, and G. Ning, "Class Weights Random Forest Algorithm for Processing Class Imbalanced Medical Data," IEEE Access, Vol. 6, pp. 4641–4652, 2018, DOI: 10.1109/ACCESS.2018.2789428.

S. M. Lundberg and S.-I. Lee, "A Unified Approach to Interpreting Model Predictions," in Proc. 31st Int. Conf. Neural Information Processing Systems (NIPS), Long Beach, CA, USA, 2017, pp. 4768–4777.

T. Saito and M. Rehmsmeier, "The Precision-Recall Plot is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets," PLoS ONE, Vol. 10, No. 3, art. e0118432, 2015, DOI: 10.1371/journal.pone.0118432.

A. A. Noor, A. Manzoor, M. D. M. Qureshi, M. A. Qureshi, and W. Rashwan, "Unveiling Explainable AI in Healthcare: Current Trends, Challenges, and Future Directions," WIREs Data Min. Knowl. Discov., Vol. 15, No. 2, art. e70018, 2025, DOI: 10.1002/widm.70018.

T. G. Dietterich, "Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms," Neural Comput., Vol. 10, No. 7, pp. 1895–1923, 1998, DOI: 10.1162/089976698300017197.

N. Chandrasekhar and S. Peddakrishna, "Enhancing Heart Disease Prediction Accuracy through Machine Learning Techniques and Optimization," Processes, Vol. 11, No. 4, art. 1210, 2023, DOI: 10.3390/pr11041210.

D. Asif, M. Bibi, M. S. Arif, and A. Mukheimer, "Enhancing Heart Disease Prediction through Ensemble Learning Techniques with Hyperparameter Optimization," Algorithms, Vol. 16, No. 6, art. 308, 2023, DOI: 10.3390/a16060308.

A. Naik, G. G. Tejani, and S. J. Mousavirad, "SGO Enhanced Random Forest and Extreme Gradient Boosting Framework for Heart Disease Prediction," SCI. Rep., Vol. 15, art. 18145, 2025, DOI: 10.1038/s41598-025-02525-7.

H. Sadr, A. Salari, M. T. Ashoobi, and M. Nazari, "Cardiovascular Disease Diagnosis: A Holistic Approach using the Integration of Machine Learning and Deep Learning Models," Eur. J. Med. Res., Vol. 29, No. 1, art. 455, 2024, DOI: 10.1186/s40001-024-02044-7.

C.-f. Chen, Z.-y. Ren, H.-h. Zong, Y.-t. Xiong, and Y. Hong, "Development and Validation of Explainable Machine Learning Models for Predicting 3-Month Functional Outcomes in Acute Ischemic Stroke: A SHAP-based Approach," Front. Neurol., Vol. 16, art. 1678815, 2025, DOI: 10.3389/fneur.2025.1678815.

A. Ogunpola, F. Saeed, S. Basurra, A. M. Albarrak, and S. N. Qasem, "Machine Learning-based Predictive Models for Detection of Cardiovascular Diseases," Diagnostics, Vol. 14, No. 2, art. 144, 2024, DOI: 10.3390/diagnostics14020144.

B. Chulde-Fernández, D. Enríquez-Ortega, C. Guevara, P. Navas, A. Tirado-Espín, P. Vizcaíno-Imacaña, F. Villalba-Meneses, C. Cadena-Morejón, D. Almeida-Galárraga, and P. Acosta-Vargas, "Classification of Heart Failure using Machine Learning: A Comparative Study," Life, Vol. 15, No. 3, art. 496, 2025, DOI: 10.3390/life15030496.

N. Selayanti, S. A. Putri, M. Kristanaya, M. P. Azzahra, M. G. Navsih, and K. M. Hindrayani, "Penerapan Machine Learning Algoritma Random Forest untuk Prediksi Penyakit Jantung," Prosiding Semin. Nas. Sains Data, Vol. 4, No. 1, pp. 895–906, 2024, DOI: 10.33005/senada.v4i1.376.

C. Elkan, "The Foundations of Cost-Sensitive Learning," in Proc. 17th Int. Joint Conf. Artificial Intelligence (IJCAI), Seattle, WA, USA, 2001, Vol. 2, pp. 973–978.

A. Fernández, S. García, M. Galar, R. C. Prati, B. Krawczyk, and F. Herrera, Learning from Imbalanced Data Sets. Cham, Switzerland: Springer, 2018, DOI: 10.1007/978-3-319-98074-4.

A. J. Vickers and E. B. Elkin, "Decision Curve Analysis: A Novel Method for Evaluating Prediction Models," Med. Decis. Making, Vol. 26, No. 6, pp. 565–574, 2006, DOI: 10.1177/0272989X06295361.

T. Chen and C. Guestrin, "XGBoost: A Scalable Tree Boosting System," in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, San Francisco, CA, USA, 2016, pp. 785–794, DOI: 10.1145/2939672.2939785.




DOI: https://doi.org/10.32520/stmsi.v15i9.6482

Article Metrics

Abstract view : 0 times
PDF - 0 times

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.