A systematic machine learning framework integrating data preprocessing, class-imbalance handling, hyperparameter optimization, and five-fold cross-validation for heart attack risk prediction.
Cardiovascular diseases and heart attacks remain major healthcare challenges worldwide. Early identification of individuals at risk can support timely intervention and improved healthcare decision-making.
This research develops a comparative machine learning framework for heart attack risk prediction by combining systematic data preprocessing, feature preparation, class-imbalance handling, hyperparameter optimization, and cross-validation.
Multiple supervised machine learning algorithms are evaluated to determine their effectiveness in predicting heart attack risk and identifying high-risk cases.
The study emphasizes that accuracy alone is not sufficient for evaluating healthcare prediction models, particularly when correctly identifying positive/risk cases is clinically important.
The main objectives of this research are to:
- 🫀 Develop a machine learning framework for heart attack risk prediction
- 🧹 Apply systematic data preprocessing
- 🔍 Identify relevant predictive features
- ⚖️ Address class imbalance using SMOTE and class weighting
- 🤖 Compare multiple machine learning algorithms
- ⚙️ Optimize model hyperparameters
- 🔄 Apply five-fold cross-validation
- 📊 Evaluate models using multiple performance metrics
- 🎯 Investigate minority-class/heart-attack detection
- 🏆 Identify models with the most useful predictive characteristics
The research uses a Heart Attack Risk Prediction Dataset obtained from Kaggle.
The target variable is:
Heart Attack Risk
It is treated as a binary classification problem:
1 → Heart Attack Risk
0 → No Heart Attack Risk
The study considers demographic, physiological, lifestyle, and clinical-related variables. Examples include:
- Age
- Sex
- Cholesterol
- Systolic blood pressure
- Diastolic blood pressure
- Smoking
- Diabetes
- Family history
- Obesity
- Alcohol consumption
- Exercise
- BMI
- Stress level
- Previous heart problems
- Physical activity
- Sleep
- Other patient-related characteristics
The research methodology describes the target as a binary dependent variable and identifies age, sex, cholesterol, blood pressure, smoking, and diabetes among the predictor variables.
The study follows a quantitative machine-learning research methodology.
🫀 HEART ATTACK DATA
│
▼
🧹 DATA PREPROCESSING
│
┌─────────────┴─────────────┐
│ │
▼ ▼
Missing Values Data Cleaning
Duplicates Outliers
Inconsistencies Feature Preparation
│ │
└─────────────┬─────────────┘
▼
🔍 FEATURE SELECTION
│
▼
⚖️ CLASS BALANCING
│
┌─────────┴─────────┐
▼ ▼
SMOTE Class Weighting
│ │
└─────────┬─────────┘
▼
🤖 MODEL TRAINING
│
▼
⚙️ HYPERPARAMETER
OPTIMIZATION
│
▼
🔄 5-FOLD CV
│
▼
📊 MODEL EVALUATION
│
┌───────────────┼───────────────┐
▼ ▼ ▼
Accuracy Recall F1
│ │ │
└───────────────┼───────────────┘
▼
🏆 COMPARISON
The dataset was divided into 80% training and 20% testing data, followed by five-fold cross-validation for more robust performance estimation and reduced overfitting risk.
Before model development, the dataset was examined for:
- Missing values
- Duplicate records
- Inconsistent observations
- Potential outliers
- Data quality issues
Descriptive statistics were used to understand the characteristics of the dataset before model development.
Feature selection was subsequently performed to identify variables considered most relevant to heart attack prediction.
Class imbalance is an important issue in healthcare machine learning because models may become biased toward the majority class.
The research incorporates Synthetic Minority Oversampling Technique (SMOTE) and class weighting to improve the representation and detection of heart attack cases.
The dataset contained:
2511 → No Heart Attack Risk
2511 → Heart Attack Risk
Total = 5022 observations
The original classes were therefore balanced in the analyzed dataset.
After SMOTE:
4499 → Class 0
4499 → Class 1
SMOTE increased the training representation to provide a larger balanced training set.
The research compares multiple supervised machine learning algorithms.
Used as the baseline classification model because of its simplicity, computational efficiency, and interpretability.
An ensemble learning algorithm capable of modeling nonlinear relationships.
It combines predictions from multiple decision trees to improve robustness and reduce variance.
A tree-based classification model used to analyze nonlinear relationships and improve detection of heart attack cases.
An ensemble technique that sequentially builds models to improve predictive performance.
A gradient-boosting framework designed for efficient and scalable tree-based learning.
An adaptive boosting algorithm that gives greater importance to observations incorrectly classified by previous learners.
A Multi-Layer Perceptron neural network used to investigate whether nonlinear neural-network modeling can capture complex relationships between patient characteristics and heart attack risk.
The paper describes the rationale for using these models as a comparative framework, including Random Forest for nonlinear relationships, boosting algorithms for structured healthcare data, and MLP for complex nonlinear interactions.
Hyperparameter optimization was performed to identify appropriate configurations for the machine learning algorithms.
The optimization process was integrated with model validation to improve predictive performance and reduce the likelihood of selecting models based only on a single train/test split.
The final framework therefore combines:
Model
↓
Hyperparameter Optimization
↓
Cross-Validation
↓
Best Configuration
↓
Final Evaluation
The study applies 5-fold cross-validation
Fold 1 → Train / Validate
Fold 2 → Train / Validate
Fold 3 → Train / Validate
Fold 4 → Train / Validate
Fold 5 → Train / Validate
↓
Average Performance
This allows the models to be repeatedly trained and validated on different subsets of the data.
The objective is to obtain a more reliable estimate of model performance and reduce the risk of overfitting.
Multiple metrics are used because healthcare prediction requires more than simply measuring overall accuracy.
Measures the proportion of correctly classified observations.
Measures how many observations predicted as positive were actually positive.
Measures how many actual heart attack-risk cases were correctly identified.
Combines precision and recall into a single metric.
Measures the model's ability to distinguish between positive and negative classes.
Provides detailed counts of:
- True Positives
- True Negatives
- False Positives
- False Negatives
The paper explicitly evaluates accuracy, precision, recall, F1-score, ROC-AUC, balanced accuracy, and confusion matrices.
The experiments demonstrate that different models perform better under different evaluation criteria.
| Model | Approx. Accuracy | Key Observation |
|---|---|---|
| 🥇 Random Forest | ~64% | Best overall accuracy |
| 🥈 Logistic Regression | ~64% | Strong overall classification |
| 🥉 Gradient Boosting | ~63% | Strong ensemble performance |
| ⚡ LightGBM | ~62% | Strong overall performance |
| 🌳 Decision Tree | ~54% | Better minority-class detection |
| 🧠 MLP Neural Network | ~54% | Better heart-attack case detection |
The reported results show Random Forest and Logistic Regression achieving approximately 0.64 accuracy, while Gradient Boosting and LightGBM achieved above 0.62.
One of the most important findings is that the model with the highest accuracy was not necessarily the best at detecting heart attack-risk cases.
The Decision Tree and MLP Neural Network achieved higher F1-scores and recall than several other models.
Reported F1-scores were approximately:
🌳 Decision Tree → 0.37
🧠 MLP Neural Network → 0.34
⚡ LightGBM → 0.12
📈 Logistic Regression → <0.05
🌲 Random Forest → <0.05
🚀 Gradient Boosting → <0.05
This demonstrates why relying exclusively on accuracy can be misleading in healthcare prediction.
Random Forest achieved approximately:
Accuracy → 0.643
Precision → 0.494
It produced the strongest overall positive-class accuracy and precision among the compared models in the reported experiment.
The MLP demonstrated an important advantage in detecting heart attack-risk cases.
Its confusion matrix showed:
❤️ Correct Risk Cases → 215
✅ Correct Non-Risk Cases → 725
❌ Risk → Non-Risk → 413
❌ Non-Risk → Risk → 400
Compared with LightGBM, the MLP correctly detected substantially more heart attack-risk cases in the reported experiment.
This suggests a trade-off between sensitivity and specificity that is particularly important in healthcare applications.
The reported ROC-AUC values were relatively close to random classification:
🌲 Random Forest → ~0.52
⚡ LightGBM → ~0.51
🚀 Gradient Boosting → ~0.51
📈 Logistic Regression → ~0.50
🌳 Decision Tree → ~0.50
🧠 MLP → ~0.47
These results indicate that the models had limited class-separation capability on the evaluated dataset, despite differences in accuracy and F1-score.
The five-fold cross-validation results reported approximately:
Accuracy → 60%
Precision → 35%
Recall → 13%
F1-Score → 19%
This provides an additional view of model generalizability across different folds.
Random Forest and Logistic Regression achieved the strongest overall accuracy, around 64%.
Decision Tree and MLP Neural Network demonstrated stronger minority-class detection through higher recall and F1-score.
Gradient Boosting and LightGBM achieved competitive overall accuracy above 62%.
A model can achieve high overall accuracy while failing to detect a substantial number of actual heart attack-risk cases.
Five-fold cross-validation provides a more robust assessment of model generalization.
These findings support the paper's conclusion that no single model is best across every evaluation criterion.
The key contribution of this research is the development of a systematic and repeatable machine-learning framework that combines:
🧹 Data Preprocessing
+
🔍 Feature Selection
+
⚖️ Class Imbalance Handling
+
⚙️ Hyperparameter Optimization
+
🔄 Five-Fold Cross-Validation
+
📊 Multi-Metric Evaluation
↓
🫀 Heart Attack Risk Prediction
Rather than evaluating machine learning algorithms using accuracy alone, the framework emphasizes the importance of recall, precision, F1-score, ROC-AUC, and confusion matrices, especially for identifying high-risk cases.
| Technology | Purpose |
|---|---|
| 🐍 Python | Machine learning implementation |
| 🐼 Pandas | Data processing |
| 🔢 NumPy | Numerical operations |
| 🤖 Scikit-learn | Machine learning |
| ⚖️ Imbalanced-learn | SMOTE/class balancing |
| 📊 Matplotlib | Visualization |
| 🎨 Seaborn | Statistical visualization |
| 📈 Plotly | Interactive visualization |
| 📋 SPSS | Statistical analysis |
| 📓 Jupyter Notebook | Experimentation |
Heart-Attack-Risk-Prediction/
│
├── 📄 Research-Paper.docx
├── 📄 Research-Paper.pdf
├── 📓 notebooks/
│ ├── preprocessing.ipynb
│ ├── model_training.ipynb
│ └── evaluation.ipynb
│
├── 📊 results/
│ ├── confusion_matrices/
│ ├── model_comparison/
│ ├── roc_curves/
│ └── cross_validation/
│
├── 🖼️ figures/
│
└── README.md
DATASET
│
▼
🧹 PREPROCESSING
│
▼
🔍 FEATURE SELECTION
│
▼
⚖️ SMOTE / WEIGHTS
│
▼
🤖 7 ML ALGORITHMS
│
▼
⚙️ HYPERPARAMETER TUNING
│
▼
🔄 5-FOLD CV
│
▼
📊 MULTI-METRIC TESTING
│
▼
🏆 COMPARATIVE
ANALYSIS
│
▼
🫀 HEART ATTACK
RISK PREDICTION
This repository provides the research materials and supporting resources for the study:
“A Comparative Machine Learning Framework for Heart Attack Risk Prediction Using Data Preprocessing, Class Imbalance Handling, and Cross-Validation.”
The repository is intended to support:
- 🎓 Academic research
- 🔬 Reproducible experimentation
- 🤖 Machine learning research
- 🫀 Healthcare analytics
- 📊 Predictive modeling
- 📚 Further research and extension
This project is intended for academic and research purposes. The machine-learning models presented in this study are experimental predictive models and are not intended to replace professional medical diagnosis, clinical judgment, or validated clinical decision-support systems.
The reported results are based on the dataset and experimental methodology described in the research paper.
Research Area: Machine Learning • Healthcare Analytics • Cardiovascular Risk Prediction • Predictive Modeling
If you use this research or repository in your work, please cite:
Amin, A., & Azam, S. (2026).
A Comparative Machine Learning Framework for Heart Attack Risk Prediction
Using Data Preprocessing, Class Imbalance Handling, and Cross-Validation.
Data → Intelligence → Prediction → Better Decision Support