Imbalance-Aware Evaluation of Stroke Risk Prediction: A Methodological Benchmark Using Calibration and Decision Curve
Keywords: calibration, class imbalance, decision curve analysis, machine learning, stroke prediction
Abstract
Stroke prediction models are often evaluated using accuracy and ROC-AUC, although these metrics can be misleading when the outcome is rare. This study presents a methodological benchmark—not a clinical validation study—for imbalance-aware evaluation of stroke risk prediction, emphasizing probability reliability, threshold performance, and decision-analytic utility. Logistic Regression (LR), Random Forest (RF), and Gradient Boosting (GB) were evaluated using the Kaggle Stroke Prediction Dataset (n = 5,110; 249 stroke cases; prevalence = 4.87%). Training folds were balanced by random oversampling, and an analytical prior correction was applied to adjust the class-prior shift introduced by oversampling. Evaluation included ROC-AUC, precision-recall AUC, Brier score, expected calibration error, calibration slope and intercept, exploratory F2-optimal thresholds, and decision curve analysis. GB achieved the highest discrimination (AUC = 0.8415; PR-AUC = 0.2126), whereas LR showed the best calibration (ECE = 0.0048; slope = 0.967). RF produced high accuracy but very low sensitivity and severe miscalibration (slope = 0.556), making its probabilities unsuitable for threshold-based interpretation without further recalibration. LR and GB generated comparable net benefit within a prespecified exploratory threshold range of 1–10%. These findings show that calibration-aware and decision-analytic evaluation is necessary for imbalanced prediction tasks. Because the dataset has uncertain clinical provenance and no external validation was performed, the results are not intended for direct clinical implementation.
Downloads
References
Akinwumi, P. O., Ojo, S., Nathaniel, T. I., Wanliss, J., Karunwi, O., & Sulaiman, M. (2025). Evaluating machine learning models for stroke prediction based on clinical variables. Frontiers in Neurology, 16, 1668420. https://doi.org/10.3389/fneur.2025.1668420
Asadi, F., Rahimi, M., Daeechini, A. H., & Paghe, A. (2024). The most efficient machine learning algorithms in stroke prediction: A systematic review. Health Science Reports, 7(10), e70062. https://doi.org/10.1002/hsr2.70062
Biswas, N., Uddin, K. M. M., Rikta, S. T., & Dey, S. K. (2022). A comparative analysis of machine learning classifiers for stroke prediction: A predictive analytics approach. Healthcare Analytics, 2, 100116. https://doi.org/10.1016/j.health.2022.100116
Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.
Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953
Collins, G. S., Dhiman, P., Ma, J., Schlussel, M. M., Archer, L., Van Calster, B., Harrell, F. E., Martin, G. P., Moons, K. G. M., van Smeden, M., Sperrin, M., Bullock, G. S., & Riley, R. D. (2024). Evaluation of clinical prediction models (part 1): from development to external validation. BMJ, 384, e074819. https://doi.org/10.1136/bmj-2023-074819
Collins, G. S., Moons, K. G. M., Dhiman, P., Riley, R. D., Beam, A. L., Van Calster, B., Ghassemi, M., Liu, X., Reitsma, J. B., van Smeden, M., Boulesteix, A.-L., Camaradou, J. C., Celi, L. A., Denaxas, S., Denniston, A. K., Glocker, B., Golub, R. M., Harvey, H., Heinze, G., … Logullo, P. (2024). TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ, 385, e078378. https://doi.org/10.1136/bmj-2023-078378
Davis, J., & Goadrich, M. (2006). The relationship between Precision-Recall and ROC curves. Proceedings of the 23rd International Conference on Machine Learning - ICML ’06, 233–240. https://doi.org/10.1145/1143844.1143874
Dritsas, E., & Trigka, M. (2022). Stroke Risk Prediction with Machine Learning Techniques. Sensors, 22(13), 4670. https://doi.org/10.3390/s22134670
Fang, Y., Zhao, X., Gong, Z., & He, J. (2022). Comparison of machine learning methods for stroke risk prediction. Journal of Stroke and Cerebrovascular Diseases, 31.
Feigin, V. L., Krishnamurthi, R. V., Parmar, P., Norrving, B., Mensah, G. A., Bennett, D. A., Barker-Collo, S., Moran, A. E., Sacco, R. L., Truelsen, T., Davis, S., Pandian, J. D., Naghavi, M., Forouzanfar, M. H., Nguyen, G., Johnson, C. O., Vos, T., Meretoja, A., Murray, C. J. L., & Roth, G. A. (2015). Update on the Global Burden of Ischemic and Hemorrhagic Stroke in 1990-2013: The GBD 2013 Study. Neuroepidemiology, 45(3), 161–176. https://doi.org/10.1159/000441085
Fernández, A., García, S., Galar, M., Prati, R. C., Krawczyk, B., & Herrera, F. (2018). Learning from Imbalanced Data Sets. Springer International Publishing. https://doi.org/10.1007/978-3-319-98074-4
Gibson, A. D., White, N. M., Collins, G. S., & Barnett, A. G. (2026). Evidence of unreliable data and poor data provenance in clinical prediction model research and clinical practice. BMC Medicine, 24(1), 386. https://doi.org/10.1186/s12916-026-04981-y
Gorelick, P. B. (2020). Risk factors for stroke. Stroke, 51(3).
Hankey, G. J. (2017). Stroke. The Lancet, 389(10069), 641–654. https://doi.org/10.1016/S0140-6736(16)30962-X
Hassan, A., Gulzar Ahmad, S., Ullah Munir, E., Ali Khan, I., & Ramzan, N. (2024). RETRACTED ARTICLE: Predictive modelling and identification of key risk factors for stroke using machine learning. Scientific Reports, 14(1), 11498. https://doi.org/10.1038/s41598-024-61665-4
Huber, M., Schober, P., Petersen, S., & Luedi, M. M. (2023). Decision curve analysis confirms higher clinical utility of multi-domain versus single-domain prediction models in patients with open abdomen treatment for peritonitis. BMC Medical Informatics and Decision Making, 23(1), 63. https://doi.org/10.1186/s12911-023-02156-w
Johnson, J. M., & Khoshgoftaar, T. M. (2019). Survey on deep learning with class imbalance. Journal of Big Data, 6(1), 27. https://doi.org/10.1186/s40537-019-0192-5
Kementerian Kesehatan Republik Indonesia. (2018). Laporan Nasional RISKESDAS 2018.
Kerr, K. F., Brown, M. D., Zhu, K., & Janes, H. (2016). Assessing the Clinical Impact of Risk Prediction Models With Decision Curves: Guidance for Correct Interpretation and Appropriate Use. Journal of Clinical Oncology, 34(21), 2534–2540. https://doi.org/10.1200/JCO.2015.65.5654
Kohavi, R. (1995). A study of cross-validation and bootstrap for accuracy estimation and model selection. Proceedings of the 14th International Joint Conference on Artificial Intelligence, 1137–1143.
Kokkotis, C., Giarmatzis, G., Giannakou, E., Moustakidis, S., Tsatalas, T., Tsiptsios, D., Vadikolias, K., & Aggelousis, N. (2022). An Explainable Machine Learning Pipeline for Stroke Prediction on Imbalanced Data. Diagnostics, 12(10), 2392. https://doi.org/10.3390/diagnostics12102392
Lemaître, G., Nogueira, F., & Aridas, C. K. (2017). Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning. Journal of Machine Learning Research, 18(17), 1–5.
Little, R., & Rubin, D. (2019). Statistical Analysis with Missing Data, Third Edition. Wiley. https://doi.org/10.1002/9781119482260
Lundberg, S., & Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. 31st Conference on Neural Information Processing Systems (NIPS 2017), 4765–4774.
Melnykova, N., Patereha, Y., Skopivskyi, S., Farion, M., Fedushko, S., & Drohomyretska, K. (2025). Machine learning for stroke prediction using imbalanced data. Scientific Reports, 15(1), 33773. https://doi.org/10.1038/s41598-025-01855-w
Ojeda, F. M., Jansen, M. L., Thiéry, A., Blankenberg, S., Weimar, C., Schmid, M., & Ziegler, A. (2023). Calibrating machine learning approaches for probability estimation: A comprehensive comparison. Statistics in Medicine, 42(29), 5451–5478. https://doi.org/10.1002/sim.9921
Powers, D. M. W. (2020). Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation. International Journal of Machine Learning Technology, 2(1), 37–63.
Pozzolo, A. D., Caelen, O., Johnson, R. A., & Bontempi, G. (2015). Calibrating Probability with Undersampling for Unbalanced Classification. 2015 IEEE Symposium Series on Computational Intelligence, 159–166. https://doi.org/10.1109/SSCI.2015.33
Riley, R. D., Archer, L., Snell, K. I. E., Ensor, J., Dhiman, P., Martin, G. P., Bonnett, L. J., & Collins, G. S. (2024). Evaluation of clinical prediction models (part 2): how to undertake an external validation study. BMJ, 384, e074820. https://doi.org/10.1136/bmj-2023-074820
Rothwell, P., Coull, A., Silver, L., Fairhead, J., Giles, M., Lovelock, C., Redgrave, J., Bull, L., Welch, S., Cuthbertson, F., Binney, L., Gutnikov, S., Anslow, P., Banning, A., Mant, D., & Mehta, Z. (2005). Population-based study of event-rate, incidence, case fatality, and mortality for all acute vascular events in all arterial territories (Oxford Vascular Study). The Lancet, 366(9499), 1773–1783. https://doi.org/10.1016/S0140-6736(05)67702-1
Saito, T., & Rehmsmeier, M. (2015). The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLOS ONE, 10(3), e0118432. https://doi.org/10.1371/journal.pone.0118432
Shoaib, M., Ali, N., Ali, S., Khattak, A., & Islam, N. (2023). Stroke prediction using machine learning with oversampling. Computers in Biology and Medicine, 161, 107021.
Soriano, F. (2021). Healthcare dataset: Stroke prediction. https://www.kaggle.com/datasets/fedesoriano/stroke-prediction-dataset
Steyerberg, E. W. (2009). Clinical Prediction Models. Springer New York. https://doi.org/10.1007/978-0-387-77244-8
Van Calster, B., McLernon, D. J., van Smeden, M., Wynants, L., & Steyerberg, E. W. (2019). Calibration: the Achilles heel of predictive analytics. BMC Medicine, 17(1), 230. https://doi.org/10.1186/s12916-019-1466-7
Van Calster, B., Nieboer, D., Vergouwe, Y., De Cock, B., Pencina, M. J., & Steyerberg, E. W. (2016). A calibration hierarchy for risk models was defined: from utopia to empirical data. Journal of Clinical Epidemiology, 74, 167–176. https://doi.org/10.1016/j.jclinepi.2015.12.005
Vickers, A. J., & Elkin, E. B. (2006). Decision Curve Analysis: A Novel Method for Evaluating Prediction Models. Medical Decision Making, 26(6), 565–574. https://doi.org/10.1177/0272989X06295361
Copyright (c) 2026 Copyright (c) 2026. Dewi Juliah Ratnaningsih, Wisnu Aji Pamungkas

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.