End-to-end risk and customer analytics platform for retail banking. Covers credit-risk modeling (PD + IFRS 9 ECL staging), transaction fraud detection (rule-based + ML), customer-level revenue analytics, and a dashboard for the risk and commercial teams.
Built to mirror the analytics stack that risk and data-analytics teams operate inside major retail banks (Itaú, Bradesco, Nubank, Inter, JPMorgan, Citi). It demonstrates fluency across four areas that distinguish senior bank analysts from junior ones:
-
Credit risk modeling PD (Probability of Default) model + IFRS 9 ECL staging with realistic LGD assumptions per product. Stage 1 / 2 / 3 rules implemented per the standard.
-
Fraud detection Two-layer pipeline: a transparent rule-based engine for compliance audit trails, and a Random Forest ML model for patterns the rules miss.
-
Customer analytics Cross-sell opportunity scoring, revenue concentration analysis, dormancy detection — the commercial side of bank analytics.
-
Engineering rigor Vectorized data pipelines (1.27M transactions in 3 seconds), 12 pytest tests covering data integrity and model logic, GitHub Actions CI, type hints throughout.
Built on a synthetic dataset designed to mirror retail banking dynamics:
| Metric | Value |
|---|---|
| Customers | 30,000 |
| Accounts | 66,980 |
| Transactions | 1,267,466 (10,217 fraud — 0.81%) |
| Loans | 8,448 (1,694 defaulted — 20.05%) |
| Total credit exposure | R$ 743M |
| PD model ROC-AUC | 0.69 (realistic for credit-risk models) |
| Fraud model recall | 60% at 0.81% fraud prevalence |
This project uses domain-specific terms. Each is explained briefly here so the README is self-contained:
| Term | Meaning |
|---|---|
| PD | Probability of Default — the probability that a borrower will default in a defined window (12 months for Stage 1, lifetime for Stage 2 & 3). |
| LGD | Loss Given Default — the share of exposure that's not recovered after default. Mortgages have low LGD (~20%) due to collateral; credit lines have high LGD (~75%). |
| EAD | Exposure At Default — the amount outstanding when default occurs. |
| ECL | Expected Credit Loss = PD × LGD × EAD. The provision banks must hold against the loan. |
| IFRS 9 | International accounting standard that defines how banks must provision for credit losses. Replaced the old "incurred loss" model with a "forward-looking expected loss" model. |
| Stage 1 / 2 / 3 | IFRS 9 categories. Stage 1 = performing loans (12-month ECL). Stage 2 = significant increase in credit risk (lifetime ECL). Stage 3 = credit-impaired / defaulted (lifetime ECL on full exposure). |
| DPD | Days Past Due. The trigger for moving between stages (DPD ≥ 30 → Stage 2, DPD ≥ 90 → Stage 3). |
| Basel III | Global regulatory framework for bank capital. Drives the upstream PD/LGD/EAD parameters used in capital calculations. |
┌──────────────────┐
│ Raw CSVs │ (synthetic generator)
│ customers │
│ accounts │
│ transactions │
│ loans │
│ credit_bureau │
└────────┬─────────┘
│
┌─────────────────┼──────────────────┐
▼ ▼ ▼
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Credit Risk │ │ Fraud │ │ Customer │
│ ──────────── │ │ ──────────── │ │ ──────────── │
│ PD model │ │ Rule engine │ │ Cross-sell │
│ IFRS 9 ECL │ │ RF model │ │ Revenue │
│ Staging │ │ Pattern viz │ │ Dormancy │
└──────┬───────┘ └──────┬───────┘ └──────┬───────┘
│ │ │
└─────────────────┼──────────────────┘
▼
┌─────────────────────┐
│ Streamlit dashboard │
│ + executive report │
└─────────────────────┘
bank-risk-analytics/
├── .github/workflows/ # CI pipeline
├── data/
│ ├── raw/ # Generated CSVs
│ └── processed/ # Trained models + scored datasets
├── sql/
│ ├── ddl/ # Star schema DDL
│ └── analytics/ # Risk & customer analytics queries
├── src/
│ ├── etl/ # Data generation pipeline
│ ├── credit_risk/ # PD model + IFRS 9 staging
│ ├── fraud/ # Rule engine + ML fraud model
│ ├── customer/ # Cross-sell + dormancy + revenue
│ ├── monitoring/ # Model drift monitoring (placeholder)
│ └── utils/ # Visualization helpers
├── dashboards/ # Streamlit app
├── tests/ # Pytest suite (12 tests)
├── images/ # Chart previews
├── requirements.txt
├── README.md
└── LICENSE
git clone https://github.com/<your-username>/bank-risk-analytics.git
cd bank-risk-analytics
python -m venv venv && source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
# 1. Generate the dataset (~1.3M rows)
python src/etl/generate_data.py
# 2. Train the credit-risk model and run IFRS 9 staging
python src/credit_risk/pd_model.py
# 3. Train the fraud detection model (rule engine + RF)
python src/fraud/fraud_detection.py
# 4. Compute customer analytics
python src/customer/customer_analytics.py
# 5. Generate visualization charts
python src/utils/generate_charts.py
# 6. Run the test suite
pytest tests/ -vDefault rates, fraud patterns, and revenue distributions are not random. Default probability correlates with bureau score, debt-to-income, and past delinquencies. Fraud transactions follow four documented patterns (high amount, foreign geography, overnight, velocity). The Lorenz curve of customer revenue follows a realistic Pareto distribution. This makes downstream models meaningful instead of trivial.
Black-box models would score better on AUC, but credit-risk models are scrutinized by regulators and internal model-validation teams. Logistic regression gives explainable coefficients — the team can articulate exactly why a customer's PD increased. ROC-AUC of 0.69 is in the realistic range for retail PD models (production banks typically score 0.65–0.80; anything higher tends to mean leakage or overfitting).
A rule-based engine that compliance can audit, plus an ML model for patterns the rules miss. This is the production pattern at major banks — pure ML detection without an interpretable layer fails internal audit.
The first version of the transaction generator used a Python loop and didn't finish in 3 minutes. Rewriting with numpy vectorization brought the same 1.27M rows down to 3 seconds — a 60×+ speedup. Worth highlighting because senior analysts are expected to recognize and fix these bottlenecks without prompting.
The first ECL implementation could produce provisions exceeding the exposure (a violation of IFRS 9 by definition — you cannot lose more than the amount at risk). The fix is a min(raw_ecl, ead) cap. Caught by the test test_ecl_does_not_exceed_ead. This is the kind of correctness work that distinguishes a project that ships from a project that just runs.
- Online fraud scoring service (FastAPI) with sub-100ms latency target
- Model drift monitoring (PSI, KS, prediction-distribution shift)
- Backtesting framework for PD model recalibration
- Stress-testing module (macroeconomic scenarios)
- LGD model (replace fixed assumptions with a regression on collateral data)
- PIX-specific fraud signatures (relevant for Brazilian banks)
- Power BI / Tableau export of the executive dashboard
Matheus Raul Silvestresan — Data Analyst with focus on operational and commercial analytics for digital businesses. Currently expanding into financial-services analytics: credit risk, fraud detection, and customer modeling for retail banks.
- LinkedIn: matheus-raul-silvestresan
- Email: matheusraulm1@gmail.com
- Location: São Paulo, Brazil 🇧🇷
- Open to remote and hybrid opportunities (LATAM / global)
MIT — see LICENSE.
Note on data and domain expertise: All data is synthetic and generated programmatically. The credit-risk methodology (PD, IFRS 9 staging, ECL formula) follows publicly documented industry standards but is simplified for portfolio purposes — production-grade models require segment-specific calibration, regulator-validated assumptions, and significant domain expertise that goes beyond what a portfolio project demonstrates. This project shows engineering competence and conceptual fluency, not deep regulatory expertise.



