Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Bank Risk Analytics Platform

End-to-end risk and customer analytics platform for retail banking. Covers credit-risk modeling (PD + IFRS 9 ECL staging), transaction fraud detection (rule-based + ML), customer-level revenue analytics, and a dashboard for the risk and commercial teams.

Python SQL scikit--learn Streamlit CI License


What this project demonstrates

Built to mirror the analytics stack that risk and data-analytics teams operate inside major retail banks (Itaú, Bradesco, Nubank, Inter, JPMorgan, Citi). It demonstrates fluency across four areas that distinguish senior bank analysts from junior ones:

  1. Credit risk modeling PD (Probability of Default) model + IFRS 9 ECL staging with realistic LGD assumptions per product. Stage 1 / 2 / 3 rules implemented per the standard.

  2. Fraud detection Two-layer pipeline: a transparent rule-based engine for compliance audit trails, and a Random Forest ML model for patterns the rules miss.

  3. Customer analytics Cross-sell opportunity scoring, revenue concentration analysis, dormancy detection — the commercial side of bank analytics.

  4. Engineering rigor Vectorized data pipelines (1.27M transactions in 3 seconds), 12 pytest tests covering data integrity and model logic, GitHub Actions CI, type hints throughout.


Headline numbers

Built on a synthetic dataset designed to mirror retail banking dynamics:

Metric Value
Customers 30,000
Accounts 66,980
Transactions 1,267,466 (10,217 fraud — 0.81%)
Loans 8,448 (1,694 defaulted — 20.05%)
Total credit exposure R$ 743M
PD model ROC-AUC 0.69 (realistic for credit-risk models)
Fraud model recall 60% at 0.81% fraud prevalence

Glossary (for non-banking readers)

This project uses domain-specific terms. Each is explained briefly here so the README is self-contained:

Term Meaning
PD Probability of Default — the probability that a borrower will default in a defined window (12 months for Stage 1, lifetime for Stage 2 & 3).
LGD Loss Given Default — the share of exposure that's not recovered after default. Mortgages have low LGD (~20%) due to collateral; credit lines have high LGD (~75%).
EAD Exposure At Default — the amount outstanding when default occurs.
ECL Expected Credit Loss = PD × LGD × EAD. The provision banks must hold against the loan.
IFRS 9 International accounting standard that defines how banks must provision for credit losses. Replaced the old "incurred loss" model with a "forward-looking expected loss" model.
Stage 1 / 2 / 3 IFRS 9 categories. Stage 1 = performing loans (12-month ECL). Stage 2 = significant increase in credit risk (lifetime ECL). Stage 3 = credit-impaired / defaulted (lifetime ECL on full exposure).
DPD Days Past Due. The trigger for moving between stages (DPD ≥ 30 → Stage 2, DPD ≥ 90 → Stage 3).
Basel III Global regulatory framework for bank capital. Drives the upstream PD/LGD/EAD parameters used in capital calculations.

Visual highlights

Credit portfolio — IFRS 9 staging and product default rates

Credit portfolio

Fraud patterns — fraud transactions cluster at 2-5 AM and at high amounts

Fraud patterns

Customer revenue concentration (Lorenz curve)

Revenue concentration

Bureau score and risk-band distribution

Risk distribution


Architecture

                 ┌──────────────────┐
                 │   Raw CSVs       │   (synthetic generator)
                 │  customers       │
                 │  accounts        │
                 │  transactions    │
                 │  loans           │
                 │  credit_bureau   │
                 └────────┬─────────┘
                          │
        ┌─────────────────┼──────────────────┐
        ▼                 ▼                  ▼
 ┌──────────────┐  ┌──────────────┐   ┌──────────────┐
 │ Credit Risk  │  │   Fraud      │   │   Customer   │
 │ ──────────── │  │ ──────────── │   │ ──────────── │
 │ PD model     │  │ Rule engine  │   │ Cross-sell   │
 │ IFRS 9 ECL   │  │ RF model     │   │ Revenue      │
 │ Staging      │  │ Pattern viz  │   │ Dormancy     │
 └──────┬───────┘  └──────┬───────┘   └──────┬───────┘
        │                 │                  │
        └─────────────────┼──────────────────┘
                          ▼
                ┌─────────────────────┐
                │ Streamlit dashboard │
                │ + executive report  │
                └─────────────────────┘

Project structure

bank-risk-analytics/
├── .github/workflows/         # CI pipeline
├── data/
│   ├── raw/                   # Generated CSVs
│   └── processed/             # Trained models + scored datasets
├── sql/
│   ├── ddl/                   # Star schema DDL
│   └── analytics/             # Risk & customer analytics queries
├── src/
│   ├── etl/                   # Data generation pipeline
│   ├── credit_risk/           # PD model + IFRS 9 staging
│   ├── fraud/                 # Rule engine + ML fraud model
│   ├── customer/              # Cross-sell + dormancy + revenue
│   ├── monitoring/            # Model drift monitoring (placeholder)
│   └── utils/                 # Visualization helpers
├── dashboards/                # Streamlit app
├── tests/                     # Pytest suite (12 tests)
├── images/                    # Chart previews
├── requirements.txt
├── README.md
└── LICENSE

Quick start

git clone https://github.com/<your-username>/bank-risk-analytics.git
cd bank-risk-analytics

python -m venv venv && source venv/bin/activate     # Windows: venv\Scripts\activate
pip install -r requirements.txt

# 1. Generate the dataset (~1.3M rows)
python src/etl/generate_data.py

# 2. Train the credit-risk model and run IFRS 9 staging
python src/credit_risk/pd_model.py

# 3. Train the fraud detection model (rule engine + RF)
python src/fraud/fraud_detection.py

# 4. Compute customer analytics
python src/customer/customer_analytics.py

# 5. Generate visualization charts
python src/utils/generate_charts.py

# 6. Run the test suite
pytest tests/ -v

Key engineering decisions

Realistic correlations in synthetic data

Default rates, fraud patterns, and revenue distributions are not random. Default probability correlates with bureau score, debt-to-income, and past delinquencies. Fraud transactions follow four documented patterns (high amount, foreign geography, overnight, velocity). The Lorenz curve of customer revenue follows a realistic Pareto distribution. This makes downstream models meaningful instead of trivial.

Logistic regression for the PD model

Black-box models would score better on AUC, but credit-risk models are scrutinized by regulators and internal model-validation teams. Logistic regression gives explainable coefficients — the team can articulate exactly why a customer's PD increased. ROC-AUC of 0.69 is in the realistic range for retail PD models (production banks typically score 0.65–0.80; anything higher tends to mean leakage or overfitting).

Two-layer fraud detection

A rule-based engine that compliance can audit, plus an ML model for patterns the rules miss. This is the production pattern at major banks — pure ML detection without an interpretable layer fails internal audit.

Vectorized data generation

The first version of the transaction generator used a Python loop and didn't finish in 3 minutes. Rewriting with numpy vectorization brought the same 1.27M rows down to 3 seconds — a 60×+ speedup. Worth highlighting because senior analysts are expected to recognize and fix these bottlenecks without prompting.

IFRS 9 ECL cap at EAD

The first ECL implementation could produce provisions exceeding the exposure (a violation of IFRS 9 by definition — you cannot lose more than the amount at risk). The fix is a min(raw_ecl, ead) cap. Caught by the test test_ecl_does_not_exceed_ead. This is the kind of correctness work that distinguishes a project that ships from a project that just runs.


Roadmap

  • Online fraud scoring service (FastAPI) with sub-100ms latency target
  • Model drift monitoring (PSI, KS, prediction-distribution shift)
  • Backtesting framework for PD model recalibration
  • Stress-testing module (macroeconomic scenarios)
  • LGD model (replace fixed assumptions with a regression on collateral data)
  • PIX-specific fraud signatures (relevant for Brazilian banks)
  • Power BI / Tableau export of the executive dashboard

About the author

Matheus Raul Silvestresan — Data Analyst with focus on operational and commercial analytics for digital businesses. Currently expanding into financial-services analytics: credit risk, fraud detection, and customer modeling for retail banks.


License

MIT — see LICENSE.

Note on data and domain expertise: All data is synthetic and generated programmatically. The credit-risk methodology (PD, IFRS 9 staging, ECL formula) follows publicly documented industry standards but is simplified for portfolio purposes — production-grade models require segment-specific calibration, regulator-validated assumptions, and significant domain expertise that goes beyond what a portfolio project demonstrates. This project shows engineering competence and conceptual fluency, not deep regulatory expertise.

About

End-to-end risk analytics platform for retail banking: PD model, IFRS 9 ECL staging, fraud detection (rules + ML), and customer analytics. Built with Python and scikit-learn.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages