Comparison of Decision Tree and CatBoost for Stroke Risk Prediction Using the Healthcare Stroke Prediction Dataset: A Data Science Approach
Keywords:
CatBoost; Data Science; Decision Tree; Machine Learning; Stroke PredictionAbstract
Stroke poses a significant health risk, necessitating an approach that can effectively predict risk based on an individual's health characteristics. This study falls within the field of Data Science specifically machine learning-based classification and aims to compare the performance of Decision Tree and CatBoost algorithms in predicting stroke risk using the Healthcare Stroke Prediction Dataset. This comparison was conducted because Decision Trees offer a simple, easily interpretable classification structure, whereas CatBoost employs a gradient boosting approach that potentially yields superior predictive capabilities. The dataset comprises 5,110 records, and the research methodology included data preprocessing, an 80:20 split for training and testing data, handling class imbalance via SMOTE, and hyperparameter optimization using GridSearchCV. Performance was evaluated using Accuracy, Precision, Recall, F1-Score, and ROC-AUC. The results indicate that the Decision Tree achieved an Accuracy of 92.37%, Precision of 18.18%, Recall of 16.00%, F1-Score of 17.02%, and ROC-AUC of 62.85%, while CatBoost achieved an Accuracy of 94.03%, Precision of 17.65%, Recall of 6.00%, F1-Score of 8.96%, and ROC-AUC of 78.39%. These findings demonstrate that CatBoost outperformed the Decision Tree in terms of Accuracy and ROC-AUC, whereas the Decision Tree performed better regarding Precision, Recall, and F1-Score. Thus, both algorithms exhibit distinct advantages in stroke risk prediction.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Rahma Yuni Simanullang (Author); Zulham Sitorus

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.










