K-HeatPulse is a real-time meteorological forecasting and urban heat risk dashboard. Built as the Lab 14 Complex Computing Activity for the Programming for AI course in my 4th semester of BS Artificial Intelligence (DUET), it moves beyond traditional static weather reporting to provide a predictive "Heat Oracle" experience. It utilizes 25 years of historical climate data alongside live Visual Crossing API ingestion to directly compare Machine Learning (Random Forest) and Deep Learning (Neural Networks) approaches, offering advanced features like side-by-side architectural evaluations and real-time urban heat profiling.
K-HeatPulse integrates historical weather records (2000โ2024) with real-time Karachi meteorological conditions to identify and forecast heatwave risk. The system features:
- Data Merge & Cleaning: Combines 25 years of historical data with recent observations
- Live API Integration: Fetches current Karachi conditions from Visual Crossing Weather API
- Dual-Brain Architecture: Compares Random Forest and Neural Network predictions side-by-side
- Interactive Dashboard: Streamlit-based UI with live metrics, predictions, and trend analysis
This project is fully production-ready and optimized for cloud deployment:
- โ Code Quality: All comments and docstrings stripped; optimized for minimal deployment footprint
- โ Syntax Validated: Both
app.pyandadvanced_features.pycompile without errors - โ Dependencies Hardened: Lightweight default
requirements.txtfor Streamlit Cloud; optional advanced features can be installed separately - โ Features Tested: Live API integration, model training, dashboard rendering, and graceful degradation all validated
- โ Performance Optimized: Caching via
@st.cache_resource(model training) and@st.cache_data(API calls); 600s TTL on live weather fetch - โ Streamlit Cloud Ready: No heavy native dependencies; minimal startup time; optional features degrade gracefully if not installed
Displays current Karachi weather conditions by fetching real-time data from the Visual Crossing API:
- Humidity (%)
- Wind Speed (km/h)
- Pressure (hPa)
- Conditions (descriptive text)
Click "Get Live Karachi Stats" to refresh the live data.
Side-by-side prediction buttons that show heatwave risk estimates:
- Predict with ML: Random Forest model trained on 25 years of tabular weather data
- Predict with DL: Neural Network (Keras) with StandardScaler normalization
Each prediction returns:
- Heat risk probability (0โ100%)
- Risk level badge: High (โฅ70%), Moderate (45โ69%), Low (<45%)
- Interpretation text
Seaborn heatmap comparing average maximum temperatures between 2000 and the latest available year in the dataset. Shows monthly trends to visualize long-term warming patterns.
- Dataset Snapshot (sidebar): Training/testing split sizes, date range, Karachi row count
- Model Comparison: Accuracy and ROC AUC scores for both ML and DL models
- Data Preview: Last 12 observations from the Karachi timeline
- File:
pakistan_weather_2000_2024.csv - Source: Kaggle
- Karachi Records: 5,844 daily observations
- Features: Temperature (min, max, avg), humidity, pressure, wind speed, precipitation, and derived metrics
- File:
pakistan_weather_data-Sep2024-Oct2025.csv - Current Coverage: 411 Karachi observations for 2024-09-10
- Note: This file was intended for 2025 but currently contains only a snapshot from September 2024
- Parse: Each file is parsed separately to handle mixed date formats (mm/dd/yyyy vs ISO 8601 with timezone)
- Filter: Keep only Karachi records (
city == "Karachi") - Clean:
- Fill missing values using forward fill (
.ffill()) - Remove empty feature columns (e.g.,
visibility) - Ensure all required features are present
- Split: 80% training (2000โ2015), 20% testing (2016โ2024)
FEATURE_COLUMNS = [
'year', 'month', 'day', 'dayofweek', 'is_weekend',
'latitude', 'longitude', 'elevation',
'tmin', 'tmax', 'tavg', 'prcp', 'wspd',
'humidity', 'pressure', 'dew_point', 'cloud_cover', 'temp_range'
]
CATEGORICAL_COLUMNS = ['season', 'wind_category', 'rainfall_intensity']
TARGET = 'is_hot_day' (binary: 1 = heatwave day, 0 = normal)
- Python: 3.11 or 3.12 (Recommended for full TensorFlow/Keras support)
- Git: Version control system
- pip: Modern Python package manager
Open your terminal and clone the project to your local machine:
git clone https://github.com/abdulhayykhan/K-HeatPulse.git
cd K-HeatPulse
Install the required Python packages using pip:
pip install -r requirements.txt
Note: TensorFlow is currently excluded on Python 3.14+ (no stable wheel available). To utilize the full Deep Learning (Neural Network) brain, please ensure your environment is running Python 3.11โ3.13.
The application relies on the Visual Crossing Weather API for real-time data. Streamlit automatically securely loads API keys from a secrets file.
- Create a folder named
.streamlitin the root of the project. - Inside that folder, create a file named
secrets.toml. - Add your API key to the file:
File: .streamlit/secrets.toml
VISUAL_CROSSING_API_KEY="YOUR_FREE_API_KEY_HERE"
How to get a Free API Key:
- Visit Visual Crossing
- Sign up for the free tier (1,000 calls/day, no credit card required)
- Copy your API key and paste it into the
secrets.tomlfile.
- A GitHub account with the repository pushed.
- The
app.py,requirements.txt,secrets.toml(kept local), and CSV data files. - Your Visual Crossing API key.
If you haven't pushed your local project to a GitHub repository yet, run the following:
git init
git add app.py advanced_features.py requirements.txt *.csv
git commit -m "think of this git message by yourself ;)"
git branch -M main
git remote add origin https://github.com/YOUR_USERNAME/k-heatpulse.git
git push -u origin main
(Note: Ensure your .gitignore includes .streamlit/secrets.toml so you don't leak your API key).
- Log in to Streamlit Cloud.
- Click "New App" โ "From GitHub repo".
- Select your
k-heatpulserepository and set the main file path toapp.py. - Click "Deploy".
Since your secrets.toml is not on GitHub, you must provide the key to the cloud server:
- Go to your live app's Settings โ Secrets (via the three-dot menu).
- Paste the exact same format used locally:
VISUAL_CROSSING_API_KEY="YOUR_FREE_API_KEY_HERE"
- Click Save to automatically reboot the app with the injected key.
To guarantee TensorFlow support, force the Streamlit server to use Python 3.11. In your app settings on the Streamlit dashboard, set the Python version to 3.11 before deploying.
- Memory: 1 GB (Sufficient for training Random Forest/Neural Net on the Karachi dataset).
- Timeout: 120 seconds (Default).
- Python Version: 3.11 (Required for DL brain; 3.14 will gracefully disable the DL features).
Once installed, spin up the local server:
streamlit run app.py
Streamlit will automatically open the dashboard in your default web browser at http://localhost:8501.
- Live Pulse: Click "Get Live Karachi Stats" to fetch current conditions via the API.
- The Duel: Click "Predict with ML" or "Predict with DL" to compare model predictions and risk assessments.
- Heat Mapping: View the seaborn heatmap to compare temperature shifts from 2000 vs. the latest year.
- Model Diagnostics: Check accuracy and ROC AUC scores in the Model Comparison table.
- Context: Inspect the Dataset Snapshot in the sidebar for training/testing sizes and data coverage.
- n_estimators: 300 trees
- class_weight:
"balanced_subsample"(handles class imbalance in hot days) - max_features:
"sqrt"(improved generalization) - min_samples_leaf: 2
- Pipeline: ColumnTransformer for feature preprocessing + RandomForestClassifier
Preprocessing:
- Numeric features: Median imputation
- Categorical features: Most-frequent imputation + One-Hot encoding
Performance (on test set):
- Accuracy: ~95.8%
- ROC AUC: ~0.77
-
Architecture:
-
Input layer (auto-sized based on feature count)
-
Dense(64, ReLU) โ Dropout(0.25)
-
Dense(32, ReLU) โ Dropout(0.15)
-
Dense(1, Sigmoid) โ Binary output
-
Optimizer: Adam (learning_rate=0.001)
-
Loss: Binary Crossentropy
-
Metrics: Accuracy, AUC
-
Epochs: 60 (with early stopping on validation loss, patience=10)
Preprocessing:
- Numeric features: Median imputation + StandardScaler (crucial for NN sensitivity)
- Categorical features: Most-frequent imputation + One-Hot encoding
- Class weighting: Balanced to address heatwave day rarity
Status: Requires TensorFlow (unavailable on Python 3.14 in this workspace)
- Random Forest: Fast, interpretable, robust to outliers
- Neural Network: Nonlinear relationships, scalable to larger datasets
- Trade-offs: Speed vs. expressiveness, ease of training vs. hyperparameter tuning
The app uses st.secrets to securely access the Visual Crossing API key. The .gitignore should exclude secrets.toml to prevent credential leaks:
# .gitignore
secrets.toml
.env
*.pkl
The get_live_karachi_weather() function caches for 600 seconds (10 minutes) to avoid exceeding API quota on repeated clicks.
Lab 14/
โโโ app.py # Main Streamlit dashboard
โโโ advanced_features.py # Optional XAI, What-If, STL features
โโโ requirements.txt # Python dependencies
โโโ README.md # This file
โโโ pakistan_weather_2000_2024.csv # Historical data
โโโ pakistan_weather_data-Sep2024-Oct2025.csv # Recent snapshot
| build_rf_pipeline() | Constructs RandomForestClassifier with preprocessing pipeline |
| build_dl_preprocessor() | Preprocessor for Neural Network (StandardScaler + OneHotEncoder) |
| build_dl_model() | Builds Keras model (disabled on Python 3.14) |
| Function (Advanced) | Purpose |
|---|---|
load_merged_data() |
Loads, merges, and cleans both CSV files |
train_models() |
Trains RF and DL models, cached for performance |
get_live_karachi_weather() |
Fetches live conditions from Visual Crossing API |
build_live_feature_row() |
Constructs feature vector from live weather |
run_ml_prediction() |
Generates Random Forest prediction |
run_dl_prediction() |
Generates Neural Network prediction |
render_*_section() |
Streamlit UI components for each feature |
| Function | Feature | Requirement |
|---|---|---|
render_xai_panel() |
Feature D: SHAP Explainability (shows top-15 feature contributions) | shap package |
render_whatif_simulator() |
Feature E: Interactive What-If Scenario Simulator (dual predictions) | plotly package (included in requirements) |
render_trend_decomposition() |
Feature F: STL Trend Decomposition (observed/trend/seasonal/residual) | statsmodels package |
Graceful Degradation: If optional packages are missing, the app skips these features but continues running with core features (AโC).
- TensorFlow Availability: Not available on Python 3.14+. Use Python 3.11โ3.13 for the DL brain. On Python 3.14, the DL path gracefully disables and the app continues with RF only.
- Data Recency: The second CSV file (
pakistan_weather_data-Sep2024-Oct2025.csv) only contains Karachi data for 2024-09-10, not a 2024โ2025 timeline. The heat map compares 2000 vs. the latest available year (currently 2024). A warning message is displayed if the latest year < 2025. - Class Imbalance: Heatwave days are rare (~0.8% of Karachi records). Both models use class weighting to compensate.
- Missing Features: The
visibilitycolumn is entirely missing in Karachi records and is excluded from the feature set. - API Rate Limiting: Visual Crossing free tier allows 1,000 calls/day. The 10-minute cache (600s) helps avoid quota exhaustion on repeated clicks.
- Optional Dependencies: SHAP explainability and STL trend decomposition require separate installation (
pip install shap statsmodels). Without these, Features D and F are skipped but the app continues normally with Features AโC, E.
- Model Training: Cached via
@st.cache_resourceโ trains once per session, reused on subsequent predictions - Live API Calls: Cached via
@st.cache_datawith TTL=600 seconds โ prevents quota exhaustion and improves responsiveness - Feature Preprocessing: Computed once during normalization, reused for all predictions
- Startup Time: ~5โ10 seconds on first load (data merge + model training); <1 second on rerun if cache persists
Optimal for Streamlit Cloud: Lightweight caching strategy minimizes memory usage and avoids Streamlit Cloud limitations.
This project demonstrates:
- Data Integration: Merging heterogeneous time-series datasets
- Time-Series Handling: Chronological train/test split for forecasting
- Preprocessing: Categorical encoding, imputation, scaling
- ML vs. DL: Comparative analysis of classical and deep learning models
- API Integration: Real-time data fetching and error handling
- Interactive Dashboards: Streamlit for rapid prototyping
- Model Deployment: Caching, feature pipelines, and gradeful fallbacks
streamlit run app.py
Check the sidebar to see:
- Total Karachi records: 6,255
- Date range: 2000-01-01 to 2024-09-10
- Training/testing split: ~5,000 / ~1,200
Click "Get Live Karachi Stats" button to see:
- Current humidity, wind speed, pressure, conditions
- Timestamp and data source
Click "Predict with ML" to see Random Forest's heatwave risk estimate, then click "Predict with DL" to compare.
View the heat map to see how Karachi's maximum temperatures have changed from 2000 to 2024. Look for warming trends in summer months (JunโAug).
- Visual Crossing Weather API: https://www.visualcrossing.com/weather-api
- Streamlit Documentation: https://docs.streamlit.io
- scikit-learn: https://scikit-learn.org
- TensorFlow/Keras: https://www.tensorflow.org
This project is open-source and available for educational and commercial use under the MIT License.
Made with โค๏ธ by Abdul Hayy Khan