End-to-end classification project that predicts the probability of loan default (loan_default) using a 20,000-record dataset of personal and financial information of approved loan applicants.
The project covers the full data-science workflow — data cleaning, null-value imputation, outlier handling, feature engineering, encoding, scaling, multiple model families, hyperparameter tuning, a neural-network comparison, and SHAP-based explainability — and concludes with a model selection that prioritises recall because, in lending, missing a defaulter (false negative) is more costly than rejecting a good borrower.
Originally developed as the second project for the Data Analytics course of the Master's in Business Analytics, ML and AI at Universidad Francisco de Vitoria (UFV).
Given an approved loan and the borrower's profile (demographics, income, credit history, debt ratios, etc.), predict whether the borrower will default.
- Target:
loan_default(binary — 1 = default, 0 = paid) - Class imbalance: only 4.76% of records correspond to defaults
- Forbidden features:
LoanApprovedandRiskScorewere excluded from training because they leak the target
| Property | Value |
|---|---|
| Rows | 20,000 (19,823 after outlier handling) |
| Columns | 35 (5 categorical, 1 date, 3 binary, 26 numeric) |
| Target | loan_default |
| Default rate | 4.76% |
The cleaned dataset is included as data/loan_clean.csv. A description of every column is provided in data/README.md.
- Data understanding — type inspection, decimal-point cleanup, date parsing.
- Missing-value treatment — domain-driven imputations:
AnnualIncome↔MonthlyIncome(×12 / ÷12)MonthlyLoanPayment≈ (LoanAmount×InterestRate) /LoanDurationMonthlyDebtPayments=DebtToIncomeRatio×MonthlyIncomeNetWorth=TotalAssets−TotalLiabilitiesCreditScoreimputed via linear regression on related credit features- Median imputation for the rest
- Outlier handling — IQR with a 3× factor; rows dropped when outliers < 1% of the column, otherwise winsorised at the 1st / 99th percentile.
- Univariate, bivariate, multivariate and temporal analysis — distributions, correlation heatmap, per-category default rates, monthly / weekday / yearly default patterns.
- Feature selection — drop redundant pairs (
Experience,MonthlyIncome,BaseInterestRate,NetWorth) and the rawApplicationDateafter decomposition. - Train / test split — 80 / 20, stratified, with a parallel "distance-model" split that also gets numeric features standardised.
- Encoding & scaling
OneHotEncoderfor nominal categoricalsOrdinalEncoderforEducationLevelStandardScalerfor numeric features (only for distance-based models)- Sin / cos transformation for cyclic temporal features (
Month,Day,DayOfWeek)
- Modelling — every model trained with
class_weight="balanced"(orscale_pos_weightfor XGBoost) to tackle the class imbalance, then tuned withGridSearchCV(cross-validated). - Neural network — PyTorch MLP (3 hidden layers, dropout 0.3, 300 epochs, lr 1e-3) for comparison.
- Explainability — coefficients of the selected linear model + SHAP summary plots.
Performance on the held-out test set (recall = sensitivity to defaulters, the metric we optimise for):
| Model | Recall | Precision | F1 | AUC-ROC | Notes |
|---|---|---|---|---|---|
| Baseline Logistic (no balancing) | 0.00 | 0.00 | 0.00 | ~0.65 | predicts "everyone pays" |
| Logistic Regression (tuned) | 0.59 | 0.07 | 0.13 | ~0.65 | best recall after Decision Tree |
| Ridge Classifier (selected) | 0.5745 | 0.0785 | 0.1381 | 0.6488 | best F1, regularised → robust |
| Decision Tree | 0.63 | 0.06 | 0.11 | 0.6 | risk of overfitting |
| Random Forest | 0.63 | 0.06 | 0.11 | 0.6 | predicts almost no defaults |
| XGBoost | 0.62 | 0.05 | 0.10 | 0.6 | same issue, dominated by majority class |
| SVM | 0.59 | 0.07 | 0.13 | 0.6466 | much slower to train |
| MLP (PyTorch) | 0.49 | 0.05 | ~0.1 | — | imbalance hurts deep model |
Why F1 / Recall and not Accuracy? Predicting "everyone pays" already gives 95.24% accuracy while detecting zero defaulters. Accuracy is misleading on imbalanced problems.
- Highest F1-score across all models (0.1381) — best precision/recall trade-off.
- Strong AUC-ROC (0.6488), competitive against tree ensembles.
- Linear → interpretable coefficients for stakeholders and regulators.
- Regularised → less prone to overfitting than a single decision tree.
Increase risk: EmploymentStatus_Unemployed, TotalDebtToIncomeRatio, PreviousLoanDefaults, low CreditScore, LoanPurpose_Debt Consolidation, long LoanDuration, high InterestRate.
Decrease risk: high TotalAssets, high AnnualIncome, good UtilityBillsPaymentHistory, lower LoanAmount, older Age, EmploymentStatus_Self-Employed.
SHAP confirms the model leans on financial-strength variables (TotalAssets, AnnualIncome, LoanAmount) much more than on demographic ones — exactly what a credit-risk team would want a lending model to prioritise.
loan-default-prediction/
├── data/
│ ├── loan_clean.csv # final cleaned dataset (~5.5 MB, 20k rows)
│ └── README.md # column dictionary
├── notebooks/
│ └── loan_default_prediction.ipynb # full end-to-end notebook
├── reports/
│ ├── project_brief_es.pdf # original course brief (Spanish)
│ ├── full_report_es.pdf # final written report (Spanish)
│ ├── full_report_es.docx # editable version of the report
│ ├── modeling_report_es.pdf # modelling deep-dive (Spanish)
│ └── eda_report_es.pdf # EDA deep-dive (Spanish)
├── figures/ # all plots used in the notebook & reports
├── src/ # (placeholder for future modular code)
├── requirements.txt
├── LICENSE
└── README.md
# 1. Clone
git clone https://github.com/AlejandroMMunizC/loan-default-prediction.git
cd loan-default-prediction
# 2. (Recommended) create a virtual environment
python -m venv .venv
source .venv/bin/activate # on Windows: .venv\Scripts\activate
# 3. Install dependencies
pip install -r requirements.txt
# 4. Launch Jupyter
jupyter lab notebooks/loan_default_prediction.ipynbThe notebook reads ../data/loan_clean.csv and runs end-to-end on a laptop in a few minutes (the SVM grid-search is the slowest step).
- Python 3.10+
- pandas, numpy — data wrangling
- matplotlib, seaborn — visualisation
- scikit-learn — preprocessing, models, metrics, GridSearchCV
- xgboost — gradient boosting
- PyTorch — neural-network baseline
- SHAP — model explainability
- Jupyter Lab / Notebook
Alejandro Magdiel Muñiz Corona Master's in Business Analytics, ML & AI — Universidad Francisco de Vitoria
- GitHub: @AlejandroMMunizC
Released under the MIT License. The dataset is included for reproducibility of the academic project; please refer to its original source for any commercial use.