Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

churn-intel-ai

Predict. Segment. Retain. ML system built on 5,630 customers achieving 99.4% AUC-ROC. Quantifies ₹7.8M revenue at risk, simulates retention campaigns with upto 6,032% ROI & delivers insights via an interactive real-time Plotly Dash dashboard.

🧠 churn-intel-ai

AI-Driven Customer Churn Prediction & Revenue Retention Analytics

Python XGBoost AUC License Status Dashboard CTA


📌 Project Overview

churn-intel-ai is an end-to-end machine learning system designed to predict customer churn in an e-commerce environment — before it happens. The system goes beyond standard classification by integrating Customer Lifetime Value (CLV) modelling, revenue-at-risk quantification, retention ROI simulation, and most importantly — per-customer personalised action intelligence that tells a business exactly who to contact, why, and what to offer them.

Built on a dataset of 5,630 customers across 20 behavioral and transactional features, the system achieves a near-perfect AUC-ROC of 99.4% using XGBoost, identifies over ₹7.8 Million in revenue at risk, and auto-generates retention strategies for the top 50 highest-risk customers based on their individual behavioral parameters.


⚙️ Tech Stack

Layer Tools
🐍 Language Python 3.11+
📊 Data Processing Pandas, NumPy
📉 Visualization Matplotlib, Seaborn, Plotly
🤖 Machine Learning Scikit-Learn, XGBoost, Imbalanced-Learn
💹 Financial Modelling Custom CLV & Revenue-at-Risk Engine
🎯 CTA Engine Rule-based Per-Customer Action Generator
📡 Dashboard Plotly Dash
🗃️ Data Format Excel (.xlsx), CSV
🔧 Environment VS Code, Jupyter Notebooks, virtualenv

🗂️ Repository Structure

churn-intel-ai/
│
├── 📁 data/
│   ├── raw/                          # Original dataset (E_Commerce_Dataset.xlsx)
│   └── processed/                    # Cleaned, engineered & scored CSVs
│
├── 📓 notebooks/
│   ├── 01_EDA.ipynb                  # Exploratory Data Analysis
│   ├── 02_preprocessing.ipynb        # Feature Engineering & Data Prep
│   ├── 03_modeling.ipynb             # ML Model Training & Evaluation
│   ├── 04_clv_analysis.ipynb         # CLV + Revenue at Risk Calculation
│   └── 05_financial_simulation.ipynb # Retention ROI Simulation
│
├── 📁 src/
│   ├── preprocess.py
│   ├── model.py
│   ├── clv.py
│   └── retention_strategy.py
│
├── 📊 dashboard/
│   └── app.py                        # Interactive Plotly Dash Dashboard
│
├── 📁 outputs/
│   ├── models/                       # Saved .pkl model files
│   └── reports/                      # Generated charts & CSV reports
│
└── 📄 requirements.txt

🔬 System Architecture

Raw Data (Excel)
      │
      ▼
┌─────────────────────┐
│  EDA & Visualization │  ──► Churn patterns, correlations, risk indicators
└─────────────────────┘
      │
      ▼
┌──────────────────────────┐
│  Preprocessing & Feature  │  ──► Null imputation, encoding,
│  Engineering              │       EngagementScore, LoyaltyTier,
└──────────────────────────┘       ComplaintRisk, HighValue flag
      │
      ▼
┌────────────────────────────────────┐
│  ML Pipeline                        │
│  ├── Logistic Regression (AUC 86.7%)│
│  ├── Random Forest    (AUC 97.9%)   │
│  └── XGBoost ✅       (AUC 99.4%)   │
└────────────────────────────────────┘
      │
      ▼
┌──────────────────────────┐
│  Churn Probability Scores │  ──► Risk Segments: High / Medium / Low
└──────────────────────────┘
      │
      ▼
┌──────────────────────────┐
│  CLV + Revenue at Risk    │  ──► Per customer financial impact
└──────────────────────────┘
      │
      ▼
┌──────────────────────────┐
│  Retention ROI Simulation │  ──► 4 campaign scenarios modelled
└──────────────────────────┘
      │
      ▼
┌────────────────────────────────────┐
│  Per-Customer CTA Intelligence      │  ──► 8 risk rules → personalised
│  Engine (NEW)                       │       action per customer
└────────────────────────────────────┘
      │
      ▼
┌──────────────────────────┐
│  Plotly Dash Dashboard    │  ──► Live interactive business intelligence
└──────────────────────────┘

📊 Key Results & Findings

🤖 Model Performance

Model Accuracy Precision Recall F1 Score AUC-ROC
Logistic Regression 80.8% 82.3% 81.9% 58.6% 86.7%
Random Forest 94.7% 82.5% 84.2% 83.3% 97.9%
XGBoost ✅ 96.5% 93.6% 85.8% 89.6% 99.4%

XGBoost selected as final production model. Confusion matrix shows only 27 missed churners out of 1,126 test samples.


🔍 Top Churn Drivers (Feature Importance)

Rank Feature Type
🥇 1 LoyaltyTier_Enc Engineered
🥈 2 Tenure Original
🥉 3 Complain Original
4 HighValue Engineered
5 ComplaintRisk Engineered
6 NumberOfDeviceRegistered Original
7 PreferedOrderCat_Enc Encoded

3 out of top 5 features were custom engineered — validating the feature engineering strategy.


📉 Critical Business Insights

⚠️  New customers (0-3 months)     →  36.2% churn rate  [CRITICAL ZONE]
✅  Customers with 24+ months       →  0.0%  churn rate  [FULLY LOYAL]
🔴  Customers who complained        →  31.7% churn rate
🟢  Customers with no complaints    →  10.9% churn rate
📱  COD payment users               →  28.8% churn rate  [HIGHEST RISK]
💳  Credit Card users               →  12.9% churn rate  [LOWEST RISK]
👤  Single marital status           →  26.7% churn rate
📍  Tier 3 city customers           →  21.4% churn rate

💰 Financial Analysis Summary

Metric Value
💼 Total CLV Portfolio ₹84.1 Million
🚨 Total Revenue at Risk ₹11.3 Million
🔴 High Risk Segment Revenue at Risk ₹7,853,629
📉 12-Month Revenue Loss (no action) ~₹10 Million
🛡️ Revenue Saved (Premium Campaign) ₹3,017,094
📈 Best Campaign ROI 6,032%
⏰ Cost of Inaction Per Month ₹850,000

🧪 Retention Scenario Simulation

Scenario Customers Targeted Revenue Saved Net Gain ROI
❌ No Intervention 0 ₹0 ₹0 0%
📧 Basic Campaign (20%) All High Risk ₹3,201,654 ₹3,041,454 1,898%
🎯 Targeted Campaign (40%) High Value + High Risk ₹2,290,698 ₹2,241,498 4,556%
👑 Premium Campaign (60%) High Value + High Risk ₹3,017,094 ₹2,967,894 6,032%

🎯 Per-Customer CTA Intelligence Engine

This is the most business-critical feature of the project. Rather than just identifying at-risk customers, the system auto-generates a personalised retention action for each of the top 50 highest-risk customers based on their individual behavioral risk parameters.

8 Risk Rules Powering the CTA Engine

Rule Trigger Condition Action Generated
🆕 New Customer Tenure ≤ 3 months Assign onboarding manager, welcome kit within 48 hours
📢 Filed Complaint Complain = 1 Priority support escalation, service recovery coupon
😞 Low Satisfaction SatisfactionScore ≤ 2 Satisfaction recovery email, free delivery for 3 orders
🔴 Critical Risk ChurnProbability ≥ 0.85 Immediate phone call, personalised 3-month retention deal
⏳ No Recent Activity DaySinceLastOrder = 0 Re-engagement push notification within 12 hours
💎 High Value at Risk BusinessSegment = High Value High Risk VIP account manager, loyalty tier upgrade
⚡ Sudden Drop Signal DaySinceLastOrder ≤ 2 + ChurnProb ≥ 0.70 Investigate last order experience proactively
🚚 High Delivery Distance WarehouseToHome ≥ 25 Free express delivery offer for next 5 orders

What the CTA Table Shows Per Customer

  • ⚠️ Risk Reasons — Every specific flag that triggered their inclusion in the list
  • ✅ Recommended Action — Exact step the retention team should take
  • Red highlighted cells — The precise column value that caused the flag
  • Sortable & Filterable — Filter by complaint = 1 to get only complainers instantly

📊 Dashboard Features

The Plotly Dash dashboard (dashboard/app.py) provides a complete business intelligence interface:

📌 Section 1 — KPI Overview

  • Total customers, high-risk count, revenue at risk, avg CLV, overall churn rate

⚡ Section 2 — CTA Recommendation Cards

  • 6 company-level action cards derived directly from model insights
  • Each card maps to a specific data finding with a concrete business action

📉 Section 3 — Analytics Visualizations

  • Churn probability distribution with decision threshold line
  • Risk segment donut chart
  • CLV vs Churn Probability scatter (color-coded by risk)
  • Revenue at Risk by segment bar chart
  • Dynamic categorical churn rate explorer
  • Financial simulation dual-axis chart (Net Gain + ROI)

🎯 Section 4 — Top 50 Priority Retention Table

  • Top 50 highest revenue-at-risk customers
  • Red cell highlighting on exact risk-triggering columns
  • Sortable and filterable live table

🧠 Section 5 — Per-Customer CTA Intelligence

  • 5 summary KPIs (callbacks needed, complaints, new customers, high value at risk, total stake)
  • Full per-customer risk reasons + recommended action table
  • Tooltip support for full action text
  • Color-coded: orange for risk reasons, green for actions

🚀 Getting Started

1️⃣ Clone the Repository

git clone https://github.com/divyansh1920/churn-intel-ai.git
cd churn-intel-ai

2️⃣ Create Virtual Environment

python -m venv .venv
.venv\Scripts\activate        # Windows
source .venv/bin/activate     # Mac/Linux

3️⃣ Install Dependencies

pip install -r requirements.txt

4️⃣ Add Dataset

Place E_Commerce_Dataset.xlsx inside:

data/raw/E_Commerce_Dataset.xlsx

5️⃣ Run Notebooks in Order

notebooks/01_EDA.ipynb
notebooks/02_preprocessing.ipynb
notebooks/03_modeling.ipynb
notebooks/04_clv_analysis.ipynb
notebooks/05_financial_simulation.ipynb

6️⃣ Launch Dashboard

cd dashboard
python app.py
# Open → http://127.0.0.1:8050

📦 Requirements

pandas
numpy
matplotlib
seaborn
scikit-learn
xgboost
imbalanced-learn
plotly
dash
jupyter
ipykernel
openpyxl

Install all at once:

pip install -r requirements.txt

📁 Output Files Generated

File Location Description
df_scored.csv data/processed/ All customers with churn probability + risk segment
df_clv.csv data/processed/ CLV + Revenue at Risk per customer
xgb_model.pkl outputs/models/ Trained XGBoost model
scaler.pkl outputs/models/ Fitted StandardScaler
imputer.pkl outputs/models/ Fitted SimpleImputer
label_encoders.pkl outputs/models/ Fitted LabelEncoders
priority_retention_list.csv outputs/reports/ Top 20 high-risk customers
financial_simulation_results.csv outputs/reports/ All 4 scenario results
model_comparison.csv outputs/reports/ Model metrics comparison
Charts 01–10 (.png) outputs/reports/ All visualization outputs

🧠 Feature Engineering — Custom Features Created

Feature Formula / Logic Purpose
EngagementScore Weighted combo of HourSpendOnApp, OrderCount, CouponUsed Overall activity score
RevenuePotential CashbackAmount × OrderCount Revenue proxy per customer
ComplaintRisk Complain == 1 AND SatisfactionScore ≤ 2 High-risk behavioral flag
LoyaltyTier Tenure binned → New / Growing / Established / Loyal Tenure-based segment
HighValue CashbackAmount > 75th percentile Binary high-value flag

📌 Dataset Information

Property Detail
Source Kaggle — E-Commerce Customer Churn Dataset
Sheet E Comm
Rows 5,630 customers
Features 20 original + 5 engineered
Target Churn (0 = Retained, 1 = Churned)
Churn Rate 16.8%
Class Imbalance Handling SMOTE (Synthetic Minority Oversampling)

👤 Author

Divyansh Jain 🔗 LinkedIn 🐙 GitHub


📄 License

This project is licensed under the MIT License — feel free to use, modify, and distribute with attribution.


Built with 🧠 machine learning, 💰 financial analytics, 🎯 per-customer intelligence, and ⚡ real business thinking.

About

Predict. Segment. Retain. ML system built on 5,630 customers achieving 99.4% AUC-ROC. Quantifies ₹7.8M revenue at risk, simulates retention campaigns with upto 6,032% ROI & delivers insights via an interactive real-time Plotly Dash dashboard.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages