Skip to content

About

A hybrid AI dashboard comparing Machine Learning and Deep Learning to predict Karachi's real-time heatwave risk using live meteorological data.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

๐Ÿ”ฅ K-HeatPulse: Karachi Heat Oracle

K-HeatPulse is a real-time meteorological forecasting and urban heat risk dashboard. Built as the Lab 14 Complex Computing Activity for the Programming for AI course in my 4th semester of BS Artificial Intelligence (DUET), it moves beyond traditional static weather reporting to provide a predictive "Heat Oracle" experience. It utilizes 25 years of historical climate data alongside live Visual Crossing API ingestion to directly compare Machine Learning (Random Forest) and Deep Learning (Neural Networks) approaches, offering advanced features like side-by-side architectural evaluations and real-time urban heat profiling.


๐ŸŽฏ Project Overview

K-HeatPulse integrates historical weather records (2000โ€“2024) with real-time Karachi meteorological conditions to identify and forecast heatwave risk. The system features:

  • Data Merge & Cleaning: Combines 25 years of historical data with recent observations
  • Live API Integration: Fetches current Karachi conditions from Visual Crossing Weather API
  • Dual-Brain Architecture: Compares Random Forest and Neural Network predictions side-by-side
  • Interactive Dashboard: Streamlit-based UI with live metrics, predictions, and trend analysis

โœ… Production-Ready Status

This project is fully production-ready and optimized for cloud deployment:

  • โœ“ Code Quality: All comments and docstrings stripped; optimized for minimal deployment footprint
  • โœ“ Syntax Validated: Both app.py and advanced_features.py compile without errors
  • โœ“ Dependencies Hardened: Lightweight default requirements.txt for Streamlit Cloud; optional advanced features can be installed separately
  • โœ“ Features Tested: Live API integration, model training, dashboard rendering, and graceful degradation all validated
  • โœ“ Performance Optimized: Caching via @st.cache_resource (model training) and @st.cache_data (API calls); 600s TTL on live weather fetch
  • โœ“ Streamlit Cloud Ready: No heavy native dependencies; minimal startup time; optional features degrade gracefully if not installed

๐Ÿ“Š Features

Feature A: Live Pulse

Displays current Karachi weather conditions by fetching real-time data from the Visual Crossing API:

  • Humidity (%)
  • Wind Speed (km/h)
  • Pressure (hPa)
  • Conditions (descriptive text)

Click "Get Live Karachi Stats" to refresh the live data.

Feature B: The Duel

Side-by-side prediction buttons that show heatwave risk estimates:

  • Predict with ML: Random Forest model trained on 25 years of tabular weather data
  • Predict with DL: Neural Network (Keras) with StandardScaler normalization

Each prediction returns:

  • Heat risk probability (0โ€“100%)
  • Risk level badge: High (โ‰ฅ70%), Moderate (45โ€“69%), Low (<45%)
  • Interpretation text

Feature C: Heat Mapping

Seaborn heatmap comparing average maximum temperatures between 2000 and the latest available year in the dataset. Shows monthly trends to visualize long-term warming patterns.

Additional Sections

  • Dataset Snapshot (sidebar): Training/testing split sizes, date range, Karachi row count
  • Model Comparison: Accuracy and ROC AUC scores for both ML and DL models
  • Data Preview: Last 12 observations from the Karachi timeline

๐Ÿ“ˆ Dataset

Historical Data (2000โ€“2024)

  • File: pakistan_weather_2000_2024.csv
  • Source: Kaggle
  • Karachi Records: 5,844 daily observations
  • Features: Temperature (min, max, avg), humidity, pressure, wind speed, precipitation, and derived metrics

Recent Data (2024)

  • File: pakistan_weather_data-Sep2024-Oct2025.csv
  • Current Coverage: 411 Karachi observations for 2024-09-10
  • Note: This file was intended for 2025 but currently contains only a snapshot from September 2024

Data Processing Pipeline

  1. Parse: Each file is parsed separately to handle mixed date formats (mm/dd/yyyy vs ISO 8601 with timezone)
  2. Filter: Keep only Karachi records (city == "Karachi")
  3. Clean:
  • Fill missing values using forward fill (.ffill())
  • Remove empty feature columns (e.g., visibility)
  • Ensure all required features are present
  1. Split: 80% training (2000โ€“2015), 20% testing (2016โ€“2024)

Features Used in Models

FEATURE_COLUMNS = [
  'year', 'month', 'day', 'dayofweek', 'is_weekend',
  'latitude', 'longitude', 'elevation',
  'tmin', 'tmax', 'tavg', 'prcp', 'wspd',
  'humidity', 'pressure', 'dew_point', 'cloud_cover', 'temp_range'
]

CATEGORICAL_COLUMNS = ['season', 'wind_category', 'rainfall_intensity']

TARGET = 'is_hot_day' (binary: 1 = heatwave day, 0 = normal)


๐Ÿš€ Installation & Setup

Prerequisites

  • Python: 3.11 or 3.12 (Recommended for full TensorFlow/Keras support)
  • Git: Version control system
  • pip: Modern Python package manager

Step 1: Clone the Repository

Open your terminal and clone the project to your local machine:

git clone https://github.com/abdulhayykhan/K-HeatPulse.git
cd K-HeatPulse

Step 2: Install Dependencies

Install the required Python packages using pip:

pip install -r requirements.txt

Note: TensorFlow is currently excluded on Python 3.14+ (no stable wheel available). To utilize the full Deep Learning (Neural Network) brain, please ensure your environment is running Python 3.11โ€“3.13.

Step 3: Configure API Key

The application relies on the Visual Crossing Weather API for real-time data. Streamlit automatically securely loads API keys from a secrets file.

  1. Create a folder named .streamlit in the root of the project.
  2. Inside that folder, create a file named secrets.toml.
  3. Add your API key to the file:

File: .streamlit/secrets.toml

VISUAL_CROSSING_API_KEY="YOUR_FREE_API_KEY_HERE"

How to get a Free API Key:

  1. Visit Visual Crossing
  2. Sign up for the free tier (1,000 calls/day, no credit card required)
  3. Copy your API key and paste it into the secrets.toml file.

๐ŸŒ Streamlit Cloud Deployment

Prerequisites for Cloud Deployment

  • A GitHub account with the repository pushed.
  • The app.py, requirements.txt, secrets.toml (kept local), and CSV data files.
  • Your Visual Crossing API key.

Deployment Steps

1. Push Code to GitHub

If you haven't pushed your local project to a GitHub repository yet, run the following:

git init
git add app.py advanced_features.py requirements.txt *.csv
git commit -m "think of this git message by yourself ;)"
git branch -M main
git remote add origin https://github.com/YOUR_USERNAME/k-heatpulse.git
git push -u origin main

(Note: Ensure your .gitignore includes .streamlit/secrets.toml so you don't leak your API key).

2. Create the Streamlit Cloud App

  1. Log in to Streamlit Cloud.
  2. Click "New App" โ†’ "From GitHub repo".
  3. Select your k-heatpulse repository and set the main file path to app.py.
  4. Click "Deploy".

3. Configure Secrets in Streamlit Cloud

Since your secrets.toml is not on GitHub, you must provide the key to the cloud server:

  1. Go to your live app's Settings โ†’ Secrets (via the three-dot menu).
  2. Paste the exact same format used locally:
VISUAL_CROSSING_API_KEY="YOUR_FREE_API_KEY_HERE"
  1. Click Save to automatically reboot the app with the injected key.

4. Configure Python Environment (Optional)

To guarantee TensorFlow support, force the Streamlit server to use Python 3.11. In your app settings on the Streamlit dashboard, set the Python version to 3.11 before deploying.

Recommended Cloud Settings

  • Memory: 1 GB (Sufficient for training Random Forest/Neural Net on the Karachi dataset).
  • Timeout: 120 seconds (Default).
  • Python Version: 3.11 (Required for DL brain; 3.14 will gracefully disable the DL features).

๐ŸŽฎ Running the App Locally

Launch the Dashboard

Once installed, spin up the local server:

streamlit run app.py

Streamlit will automatically open the dashboard in your default web browser at http://localhost:8501.

Features to Explore

  1. Live Pulse: Click "Get Live Karachi Stats" to fetch current conditions via the API.
  2. The Duel: Click "Predict with ML" or "Predict with DL" to compare model predictions and risk assessments.
  3. Heat Mapping: View the seaborn heatmap to compare temperature shifts from 2000 vs. the latest year.
  4. Model Diagnostics: Check accuracy and ROC AUC scores in the Model Comparison table.
  5. Context: Inspect the Dataset Snapshot in the sidebar for training/testing sizes and data coverage.

๐Ÿง  Model Architecture & Comparison

ML Brain: Random Forest Classifier

  • n_estimators: 300 trees
  • class_weight: "balanced_subsample" (handles class imbalance in hot days)
  • max_features: "sqrt" (improved generalization)
  • min_samples_leaf: 2
  • Pipeline: ColumnTransformer for feature preprocessing + RandomForestClassifier

Preprocessing:

  • Numeric features: Median imputation
  • Categorical features: Most-frequent imputation + One-Hot encoding

Performance (on test set):

  • Accuracy: ~95.8%
  • ROC AUC: ~0.77

DL Brain: Neural Network (Keras/TensorFlow)

  • Architecture:

  • Input layer (auto-sized based on feature count)

  • Dense(64, ReLU) โ†’ Dropout(0.25)

  • Dense(32, ReLU) โ†’ Dropout(0.15)

  • Dense(1, Sigmoid) โ†’ Binary output

  • Optimizer: Adam (learning_rate=0.001)

  • Loss: Binary Crossentropy

  • Metrics: Accuracy, AUC

  • Epochs: 60 (with early stopping on validation loss, patience=10)

Preprocessing:

  • Numeric features: Median imputation + StandardScaler (crucial for NN sensitivity)
  • Categorical features: Most-frequent imputation + One-Hot encoding
  • Class weighting: Balanced to address heatwave day rarity

Status: Requires TensorFlow (unavailable on Python 3.14 in this workspace)

Why Compare Both?

  • Random Forest: Fast, interpretable, robust to outliers
  • Neural Network: Nonlinear relationships, scalable to larger datasets
  • Trade-offs: Speed vs. expressiveness, ease of training vs. hyperparameter tuning

๐Ÿ”ง Configuration

Secrets Management

The app uses st.secrets to securely access the Visual Crossing API key. The .gitignore should exclude secrets.toml to prevent credential leaks:

# .gitignore
secrets.toml
.env
*.pkl

Caching

The get_live_karachi_weather() function caches for 600 seconds (10 minutes) to avoid exceeding API quota on repeated clicks.


๐Ÿ“ Project Structure

Lab 14/
โ”œโ”€โ”€ app.py                                # Main Streamlit dashboard 
โ”œโ”€โ”€ advanced_features.py                  # Optional XAI, What-If, STL features
โ”œโ”€โ”€ requirements.txt                      # Python dependencies
โ”œโ”€โ”€ README.md                             # This file
โ”œโ”€โ”€ pakistan_weather_2000_2024.csv        # Historical data 
โ””โ”€โ”€ pakistan_weather_data-Sep2024-Oct2025.csv  # Recent snapshot 

Key Functions in app.py

| build_rf_pipeline() | Constructs RandomForestClassifier with preprocessing pipeline | | build_dl_preprocessor() | Preprocessor for Neural Network (StandardScaler + OneHotEncoder) | | build_dl_model() | Builds Keras model (disabled on Python 3.14) |

Function (Advanced) Purpose
load_merged_data() Loads, merges, and cleans both CSV files
train_models() Trains RF and DL models, cached for performance
get_live_karachi_weather() Fetches live conditions from Visual Crossing API
build_live_feature_row() Constructs feature vector from live weather
run_ml_prediction() Generates Random Forest prediction
run_dl_prediction() Generates Neural Network prediction
render_*_section() Streamlit UI components for each feature

Advanced Features (advanced_features.py)

Function Feature Requirement
render_xai_panel() Feature D: SHAP Explainability (shows top-15 feature contributions) shap package
render_whatif_simulator() Feature E: Interactive What-If Scenario Simulator (dual predictions) plotly package (included in requirements)
render_trend_decomposition() Feature F: STL Trend Decomposition (observed/trend/seasonal/residual) statsmodels package

Graceful Degradation: If optional packages are missing, the app skips these features but continues running with core features (Aโ€“C).


โš ๏ธ Limitations & Known Issues

  1. TensorFlow Availability: Not available on Python 3.14+. Use Python 3.11โ€“3.13 for the DL brain. On Python 3.14, the DL path gracefully disables and the app continues with RF only.
  2. Data Recency: The second CSV file (pakistan_weather_data-Sep2024-Oct2025.csv) only contains Karachi data for 2024-09-10, not a 2024โ€“2025 timeline. The heat map compares 2000 vs. the latest available year (currently 2024). A warning message is displayed if the latest year < 2025.
  3. Class Imbalance: Heatwave days are rare (~0.8% of Karachi records). Both models use class weighting to compensate.
  4. Missing Features: The visibility column is entirely missing in Karachi records and is excluded from the feature set.
  5. API Rate Limiting: Visual Crossing free tier allows 1,000 calls/day. The 10-minute cache (600s) helps avoid quota exhaustion on repeated clicks.
  6. Optional Dependencies: SHAP explainability and STL trend decomposition require separate installation (pip install shap statsmodels). Without these, Features D and F are skipped but the app continues normally with Features Aโ€“C, E.

โšก Performance & Caching

  • Model Training: Cached via @st.cache_resource โ€” trains once per session, reused on subsequent predictions
  • Live API Calls: Cached via @st.cache_data with TTL=600 seconds โ€” prevents quota exhaustion and improves responsiveness
  • Feature Preprocessing: Computed once during normalization, reused for all predictions
  • Startup Time: ~5โ€“10 seconds on first load (data merge + model training); <1 second on rerun if cache persists

Optimal for Streamlit Cloud: Lightweight caching strategy minimizes memory usage and avoids Streamlit Cloud limitations.


๐ŸŽ“ Learning Objectives

This project demonstrates:

  • Data Integration: Merging heterogeneous time-series datasets
  • Time-Series Handling: Chronological train/test split for forecasting
  • Preprocessing: Categorical encoding, imputation, scaling
  • ML vs. DL: Comparative analysis of classical and deep learning models
  • API Integration: Real-time data fetching and error handling
  • Interactive Dashboards: Streamlit for rapid prototyping
  • Model Deployment: Caching, feature pipelines, and gradeful fallbacks

๐Ÿ“ Example Workflow

Step 1: Start the App

streamlit run app.py

Step 2: View Dataset Info

Check the sidebar to see:

  • Total Karachi records: 6,255
  • Date range: 2000-01-01 to 2024-09-10
  • Training/testing split: ~5,000 / ~1,200

Step 3: Get Live Weather

Click "Get Live Karachi Stats" button to see:

  • Current humidity, wind speed, pressure, conditions
  • Timestamp and data source

Step 4: Make Predictions

Click "Predict with ML" to see Random Forest's heatwave risk estimate, then click "Predict with DL" to compare.

Step 5: Analyze Trends

View the heat map to see how Karachi's maximum temperatures have changed from 2000 to 2024. Look for warming trends in summer months (Junโ€“Aug).


๐Ÿ”— Resources


๐Ÿ“„ License

This project is open-source and available for educational and commercial use under the MIT License.


Made with โค๏ธ by Abdul Hayy Khan

About

A hybrid AI dashboard comparing Machine Learning and Deep Learning to predict Karachi's real-time heatwave risk using live meteorological data.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages