Skip to content

About

Air quality forecasting for 26 Indian cities on 700K+ CPCB records — an honest benchmark: per-horizon XGBoost scored against a persistence baseline, with the target leakage that inflated earlier results found and fixed. Live Streamlit dashboard, FastAPI service, provenance-tracked data.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

36 Commits

Folders and files

Repository files navigation

India Air Quality Forecasting

700K+ records · 26 cities · 12 pollutants · 5.5 years
An honest forecasting benchmark: what daily AQI prediction can and cannot do

Live demo Tests Tests count Python License

Open the live dashboard · Case Study · Key Insights · Walkthrough


Try it first: the live dashboard runs on the bundled demo database — 29,531 real CPCB daily records across 26 cities and 78,774 hourly readings, with six trained XGBoost models. No signup, no setup.


Key Results

Forecast error (MAPE, lower is better) from a rolling backtest on held-out periods — never on data the models trained on. Persistence means simply assuming tomorrow looks like today.

City 1 day 3 days 7 days 14 days
model / persistence model / persistence model / persistence model / persistence
Bengaluru 11 / 10 17 / 15 19 / 18 21 / 19
Mumbai 13 / 13 36 / 25 33 / 27 46 / 31
Hyderabad 14 / 13 32 / 24 35 / 30 41 / 32
Delhi 15 / 16 31 / 31 37 / 36 47 / 40
Kolkata 17 / 17 39 / 27 36 / 31 37 / 38
Chennai 23 / 18 41 / 27 48 / 31 49 / 33

The headline finding is a negative one, and it is the point of the project. Persistence is a very strong baseline for daily city-level AQI. Across six cities and four horizons, gradient boosting beat it reliably only for Delhi at one to three days. Three attempts to improve on that — removing target leakage, training one model per horizon, and predicting the change rather than the level — did not overturn it.

Why an earlier version claimed 0.8–3.2%

Two engineered features leaked the target into the inputs:

  • aqi_city_zscore was (today's AQI − city mean) ÷ city std — an invertible function of the value being predicted, correlating 1.000 with it.
  • Rolling means such as aqi_roll3_mean used windows that included the current day.

The forecast also held every feature at its last observed value, returning the same number for every future day. Both are fixed, and tests/test_backtesting.py fails if either regresses.

What the honest numbers say: short-range forecasting has real signal; beyond a week, historical AQI alone is close to unpredictable, and meteorological inputs — not more trees — are what is missing.


Features

  • 6-page analytics dashboard — Executive summary, trends, pollutant drill-down, city deep-dive, data quality, ML forecasting
  • Direct per-horizon XGBoost models — 66 leakage-checked features per city, one model per forecast horizon, benchmarked against persistence and climatology
  • Data provenance — Every row tagged as real/synthetic with source tracking
  • REST API — FastAPI with /forecast/{city}, /validate/{city}, /data/freshness
  • 159 tests across 10 files, including regression tests that fail if target leakage returns
  • Runs three ways — hosted on Streamlit Cloud with zero setup, docker compose up for the full PostgreSQL stack, or a bare clone against the bundled SQLite database

Architecture

┌──────────────┐    ┌─────────────┐    ┌──────────────────┐
│  CPCB CSVs   │    │  OpenAQ API │    │  Synthetic Data  │
│  250MB, 5 f. │    │  (real-time)│    │  (fallback)      │
└──────┬───────┘    └──────┬──────┘    └────────┬─────────┘
       │                   │                    │
       └───────────────────┼────────────────────┘
                           ▼
              ┌────────────────────────┐
              │   PostgreSQL (5 tables) │
              │  city_measurements      │
              │  city_hourly           │    ← 700k+ rows, provenance-tracked
              │  station_day, stations │
              └───┬────────────────┬───┘
                  │                │
         ┌────────▼───┐    ┌──────▼────────┐
         │  Dashboards │    │  FastAPI API  │
         │  Streamlit  │    │  /forecast    │
         │  :8501      │    │  /validate    │
         └──────┬──────┘    └──────┬────────┘
                │                  │
         ┌──────▼──────────────────▼──────────┐
         │         ML Forecasting Layer        │
         │                                     │
         │   feature_engineering (66 feats)    │
         │   → ml_pipeline (time split)        │
         │   → model_training (XGB, RF, MA)    │
         │   → forecasting_service (inference) │
         └─────────────────────────────────────┘

Data flow: Raw CSVs/APIs → seed_data.py → PostgreSQL → lib/ processing → Dashboard/API/Forecast


Tech Stack

Layer Technology
Models XGBoost, scikit-learn (Random Forest), Prophet
Dashboard Streamlit + matplotlib
API FastAPI + uvicorn
Database PostgreSQL in Docker; bundled SQLite for the hosted demo (SQLAlchemy)
Data pandas, numpy
Infrastructure Docker, Docker Compose
CI/CD GitHub Actions (pytest, ruff)
Testing pytest, pytest-cov (159 tests)

Quick Start

# Docker (easiest — no local PostgreSQL needed)
git clone https://github.com/PaddyCH96/india-aqi-forecasting.git
cd india-aqi-forecasting
docker compose up --build
# Dashboard at http://localhost:8501

# Or with API:
docker compose --profile api up --build
# API docs at http://localhost:8000/docs

Data in a fresh clone

A read-only SQLite database ships with the repo at data/aqi_demo.db (9.8 MB):

Rows Source Cities
29,531 CPCB daily city aggregates (real) 26
9,870 synthetic fallback, tagged is_synthetic 6

With no AQI_DB_URL set the app uses that file, so streamlit run scripts/dashboard.py works on a fresh clone with no database server. Set AQI_DB_URL to point at PostgreSQL instead — Docker and the Render blueprint both do.

Real rows are shown by default; synthetic rows are excluded until you tick "Include synthetic data" in the sidebar, or pass ?use_synthetic=true to the API. That separation is enforced in SQL, not assumed.

The larger hourly and station-level extracts (city_hour.csv at 707,875 rows, station_hour.csv at ~4M) are not committed — they are too large for the repo. The 26-city, 700K-record figures in Key Results come from those full extracts.

Data attribution: air quality measurements are published by the Central Pollution Control Board (CPCB), Government of India.

Deploy your own (free)

Streamlit Community Cloud — free, no card, no expiry:

  1. share.streamlit.io -> sign in with GitHub
  2. New app -> this repo -> branch main
  3. Main file path: scripts/dashboard.py
  4. Deploy

No database or secrets to configure — the bundled data/aqi_demo.db is used automatically. Apps sleep after ~7 days idle and wake on the next visit.

Render — if you want PostgreSQL and the REST API, render.yaml is a Blueprint: Render dashboard -> New -> Blueprint -> this repo -> Apply. Note that Render's free PostgreSQL instances expire after 30 days.

Local Setup

python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txt
cp .env.example .env
createdb india_air_quality
python scripts/seed_data.py     # bootstraps from CSV or generates synthetic
streamlit run scripts/dashboard.py

Dashboard Pages

# Page What It Shows
1 Executive Summary National snapshot, KPI cards, city ranking
2 Historical Trends Multi-city trends, seasonal decomposition, monthly averages
3 Pollutant Drill-Down Per-pollutant distributions, 9×9 correlation matrix, diurnal patterns
4 City Deep-Dive Single-city history, year-over-year bars, pollutant summary table
5 Data Quality Missing data heatmap, completeness warnings by city
6 Forecasting 24h–336h XGBoost forecast with confidence bands + AQI alerts

Key Insights

  • Delhi is an extreme outlier — mean AQI 259.5 is 2.7× higher than the next worst city
  • PM2.5 alone predicts AQI with r=0.97 — other pollutants are largely redundant for forecasting
  • Mumbai has a monitoring crisis — 61% of daily AQI records missing, worst of 26 cities
  • Winter pollution penalty varies by geography — 2.5× in the north, 1.3× in the south
  • Data quality determines accuracy — not model choice. Better monitoring > better algorithms

Project Structure

├── lib/                 # Shared library (12 modules)
│   ├── config.py        #   Constants
│   ├── db.py            #   Parameterized SQL queries
│   ├── feature_engineering.py  # 66 features per city
│   ├── model_training.py       # 5 model trainers
│   ├── forecasting_service.py  # Inference for dashboard
│   ├── charts.py, analysis.py  # Visualization + EDA
│   └── ...                     # metrics, aqi, models, utils, logging
├── scripts/             # Runnable applications
│   ├── dashboard.py     #   6-page unified dashboard
│   ├── api.py           #   FastAPI REST API
│   ├── seed_data.py     #   Database bootstrap
│   └── ingest_hourly.py #   Hourly data pipeline
├── tests/               # 159 tests
├── docs/                # Architecture, EDA, deployment, ML eval
├── models/              # (generated) legacy single-model artefacts, not used by the forecast
├── CASE_STUDY.md        # Portfolio narrative
├── INSIGHTS.md          # Top 5 findings with evidence
└── Dockerfile + docker-compose.yml

Reading Order for Recruiters

  1. Case Study — Narrative overview of the project (10 min read)
  2. Key Insights — Five defensible findings with evidence (5 min read)
  3. Live dashboard — the running system, no setup required
  4. Code — lib/ for core logic, tests/ for test coverage

Bundled demo database: CPCB daily records 2015-01 to 2020-07 across 26 cities, hourly readings 2019-01 to 2020-07 for six cities, plus a synthetic 2020-07 to 2024-12 series tagged is_synthetic. Headline figures (700K+ records, 12 pollutants, 5.5 years) come from the full CPCB extract, which is not committed. Built with open data and open source.

About

Air quality forecasting for 26 Indian cities on 700K+ CPCB records — an honest benchmark: per-horizon XGBoost scored against a persistence baseline, with the target leakage that inflated earlier results found and fixed. Live Streamlit dashboard, FastAPI service, provenance-tracked data.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages