Predicts which customers are likely to cancel their subscription, and explains why, so a retention team can prioritize outreach before it's too late.
Telecom companies lose an estimated 15-25% of subscribers per year to churn. Even a modest improvement in early detection can save significant recurring revenue. This project builds an end-to-end pipeline — from raw customer data to a usable prediction tool — to demonstrate that.
- Cleans and engineers features from raw customer/billing data
- Trains and compares a baseline (Logistic Regression) against a tuned XGBoost model, evaluated on precision/recall/F1/ROC-AUC (not just accuracy, since churn is a heavily imbalanced problem)
- Explains individual predictions with SHAP, so the output is a decision a human can act on, not just a black-box number
- Ships as an interactive Streamlit dashboard where you can plug in a customer's details and get a churn probability + top risk factors
(Add a screenshot or GIF of the running dashboard here once you deploy it — this is the single highest-impact thing you can add to this README.)
| Model | Precision | Recall | F1 | ROC-AUC |
|---|---|---|---|---|
| Logistic Regression (baseline) | 0.42 | 0.78 | ~0.55 | ~0.80 |
| XGBoost (final) | 0.51 | 0.80 | 0.63 | 0.84 |
(Run src/train.py to regenerate these numbers — they'll differ on the
real dataset vs. the synthetic sample included here.)
churn-prediction/
├── data/ # raw + sample datasets
├── src/
│ ├── generate_sample_data.py # synthetic data for development/testing
│ ├── data_prep.py # cleaning + feature engineering
│ ├── train.py # trains baseline + final model
│ └── explain.py # SHAP-based per-customer explanations
├── app/
│ └── app.py # Streamlit dashboard
├── model/ # saved model + metadata (generated)
├── tests/
│ └── test_pipeline.py # pipeline sanity tests
├── requirements.txt
├── model_card.md # what this model does & doesn't do
└── README.md
# 1. Clone and enter the repo
git clone https://github.com/YOUR_USERNAME/churn-prediction.git
cd churn-prediction
# 2. Create a virtual environment
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
# 3. Install dependencies
pip install -r requirements.txt
# 4. Get data (pick one)
python src/generate_sample_data.py # synthetic sample, works immediately
# OR download the real dataset from Kaggle (see "Dataset" below) into data/
# 5. Train the model
python src/train.py
# 6. Launch the dashboard
streamlit run app/app.pyThis repo includes a synthetic sample generator (src/generate_sample_data.py)
so you can run everything immediately without downloading anything.
For the real dataset, download Telco Customer Churn from Kaggle:
https://www.kaggle.com/datasets/blastchar/telco-customer-churn
Place the CSV in data/telco_churn.csv and update the path in train.py.
See model_card.md for a full breakdown of what this model does and doesn't do, and where it shouldn't be trusted blindly.
MIT