Skip to content

About

End-to-end loan default classification on an imbalanced dataset (20k records, 4.76% default rate), full ML pipeline from data cleaning to SHAP-based explainability, comparing Logistic Regression, Ridge, SVM, Decision Tree, Random Forest, XGBoost, and a PyTorch MLP. UFV Master's Data Analytics project.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

Loan Default Prediction

End-to-end classification project that predicts the probability of loan default (loan_default) using a 20,000-record dataset of personal and financial information of approved loan applicants.

The project covers the full data-science workflow — data cleaning, null-value imputation, outlier handling, feature engineering, encoding, scaling, multiple model families, hyperparameter tuning, a neural-network comparison, and SHAP-based explainability — and concludes with a model selection that prioritises recall because, in lending, missing a defaulter (false negative) is more costly than rejecting a good borrower.

Originally developed as the second project for the Data Analytics course of the Master's in Business Analytics, ML and AI at Universidad Francisco de Vitoria (UFV).


Problem statement

Given an approved loan and the borrower's profile (demographics, income, credit history, debt ratios, etc.), predict whether the borrower will default.

  • Target: loan_default (binary — 1 = default, 0 = paid)
  • Class imbalance: only 4.76% of records correspond to defaults
  • Forbidden features: LoanApproved and RiskScore were excluded from training because they leak the target

Dataset

Property Value
Rows 20,000 (19,823 after outlier handling)
Columns 35 (5 categorical, 1 date, 3 binary, 26 numeric)
Target loan_default
Default rate 4.76%

The cleaned dataset is included as data/loan_clean.csv. A description of every column is provided in data/README.md.

Methodology

  1. Data understanding — type inspection, decimal-point cleanup, date parsing.
  2. Missing-value treatment — domain-driven imputations:
    • AnnualIncome ↔ MonthlyIncome (×12 / ÷12)
    • MonthlyLoanPayment ≈ (LoanAmount × InterestRate) / LoanDuration
    • MonthlyDebtPayments = DebtToIncomeRatio × MonthlyIncome
    • NetWorth = TotalAssets − TotalLiabilities
    • CreditScore imputed via linear regression on related credit features
    • Median imputation for the rest
  3. Outlier handling — IQR with a 3× factor; rows dropped when outliers < 1% of the column, otherwise winsorised at the 1st / 99th percentile.
  4. Univariate, bivariate, multivariate and temporal analysis — distributions, correlation heatmap, per-category default rates, monthly / weekday / yearly default patterns.
  5. Feature selection — drop redundant pairs (Experience, MonthlyIncome, BaseInterestRate, NetWorth) and the raw ApplicationDate after decomposition.
  6. Train / test split — 80 / 20, stratified, with a parallel "distance-model" split that also gets numeric features standardised.
  7. Encoding & scaling
    • OneHotEncoder for nominal categoricals
    • OrdinalEncoder for EducationLevel
    • StandardScaler for numeric features (only for distance-based models)
    • Sin / cos transformation for cyclic temporal features (Month, Day, DayOfWeek)
  8. Modelling — every model trained with class_weight="balanced" (or scale_pos_weight for XGBoost) to tackle the class imbalance, then tuned with GridSearchCV (cross-validated).
  9. Neural network — PyTorch MLP (3 hidden layers, dropout 0.3, 300 epochs, lr 1e-3) for comparison.
  10. Explainability — coefficients of the selected linear model + SHAP summary plots.

Results

Performance on the held-out test set (recall = sensitivity to defaulters, the metric we optimise for):

Model Recall Precision F1 AUC-ROC Notes
Baseline Logistic (no balancing) 0.00 0.00 0.00 ~0.65 predicts "everyone pays"
Logistic Regression (tuned) 0.59 0.07 0.13 ~0.65 best recall after Decision Tree
Ridge Classifier (selected) 0.5745 0.0785 0.1381 0.6488 best F1, regularised → robust
Decision Tree 0.63 0.06 0.11 0.6 risk of overfitting
Random Forest 0.63 0.06 0.11 0.6 predicts almost no defaults
XGBoost 0.62 0.05 0.10 0.6 same issue, dominated by majority class
SVM 0.59 0.07 0.13 0.6466 much slower to train
MLP (PyTorch) 0.49 0.05 ~0.1 — imbalance hurts deep model

Why F1 / Recall and not Accuracy? Predicting "everyone pays" already gives 95.24% accuracy while detecting zero defaulters. Accuracy is misleading on imbalanced problems.

Why Ridge Classifier was selected

  • Highest F1-score across all models (0.1381) — best precision/recall trade-off.
  • Strong AUC-ROC (0.6488), competitive against tree ensembles.
  • Linear → interpretable coefficients for stakeholders and regulators.
  • Regularised → less prone to overfitting than a single decision tree.

Top drivers of default risk (Ridge coefficients + SHAP)

Increase risk: EmploymentStatus_Unemployed, TotalDebtToIncomeRatio, PreviousLoanDefaults, low CreditScore, LoanPurpose_Debt Consolidation, long LoanDuration, high InterestRate.

Decrease risk: high TotalAssets, high AnnualIncome, good UtilityBillsPaymentHistory, lower LoanAmount, older Age, EmploymentStatus_Self-Employed.

SHAP confirms the model leans on financial-strength variables (TotalAssets, AnnualIncome, LoanAmount) much more than on demographic ones — exactly what a credit-risk team would want a lending model to prioritise.

Repository structure

loan-default-prediction/
├── data/
│   ├── loan_clean.csv          # final cleaned dataset (~5.5 MB, 20k rows)
│   └── README.md               # column dictionary
├── notebooks/
│   └── loan_default_prediction.ipynb   # full end-to-end notebook
├── reports/
│   ├── project_brief_es.pdf    # original course brief (Spanish)
│   ├── full_report_es.pdf      # final written report (Spanish)
│   ├── full_report_es.docx     # editable version of the report
│   ├── modeling_report_es.pdf  # modelling deep-dive (Spanish)
│   └── eda_report_es.pdf       # EDA deep-dive (Spanish)
├── figures/                    # all plots used in the notebook & reports
├── src/                        # (placeholder for future modular code)
├── requirements.txt
├── LICENSE
└── README.md

Reproducing the results

# 1. Clone
git clone https://github.com/AlejandroMMunizC/loan-default-prediction.git
cd loan-default-prediction

# 2. (Recommended) create a virtual environment
python -m venv .venv
source .venv/bin/activate          # on Windows: .venv\Scripts\activate

# 3. Install dependencies
pip install -r requirements.txt

# 4. Launch Jupyter
jupyter lab notebooks/loan_default_prediction.ipynb

The notebook reads ../data/loan_clean.csv and runs end-to-end on a laptop in a few minutes (the SVM grid-search is the slowest step).

Tech stack

  • Python 3.10+
  • pandas, numpy — data wrangling
  • matplotlib, seaborn — visualisation
  • scikit-learn — preprocessing, models, metrics, GridSearchCV
  • xgboost — gradient boosting
  • PyTorch — neural-network baseline
  • SHAP — model explainability
  • Jupyter Lab / Notebook

Author

Alejandro Magdiel Muñiz Corona Master's in Business Analytics, ML & AI — Universidad Francisco de Vitoria

License

Released under the MIT License. The dataset is included for reproducibility of the academic project; please refer to its original source for any commercial use.

About

End-to-end loan default classification on an imbalanced dataset (20k records, 4.76% default rate), full ML pipeline from data cleaning to SHAP-based explainability, comparing Logistic Regression, Ridge, SVM, Decision Tree, Random Forest, XGBoost, and a PyTorch MLP. UFV Master's Data Analytics project.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages