This project implements an end-to-end machine learning pipeline to predict weekly sales. It includes data preprocessing, feature engineering, model selection, and optimization using feature importance.
datacleaning.ipynb→ Data ingestion, merging, preprocessingmodeltraining.ipynb→ Model training, evaluation, optimizationmodels/best_model_top.pklencoder_top.pklmodel_top_features.pklraw_top_features.pkl
The dataset is built by combining:
features.csvstores.csvtrain.csv
-
Dataset Merging
- Joined on
Store,Date,IsHoliday(left join)
- Joined on
-
Missing Value Handling
- Filled
MarkDowncolumns with0 - Applied forward fill (
ffill) for:- CPI
- Unemployment (grouped by Store)
- Filled
-
Feature Engineering
- Extracted:
- Year
- Month
- Week (from Date column)
- Extracted:
-
Data Transformation
- Converted
IsHolidayto integer (0/1)
- Converted
- Chronological split (80/20)
- Ensures validation on future data (real-world scenario)
| Model | MAE | RMSE | R² |
|---|---|---|---|
| Random Forest | 1813.39 | 3764.04 | 0.9706 |
| XGBoost | 3459.77 | 5676.10 | 0.9332 |
| Linear Regression | 14542.03 | 20898.35 | 0.0950 |
Insight:
Random Forest performs the best, capturing ~97% of the variance.
Top important features identified using permutation importance:
- Dept
- Size
- Store
- CPI
- Week
- Unemployment
- Type
The model was retrained using only these features.
Final Performance:
- R² Score: 0.9711
Saved in models/ for production use:
best_model_top.pkl→ Trained modelencoder_top.pkl→ One-hot encodermodel_top_features.pkl→ Processed featuresraw_top_features.pkl→ Input features
- Python
- Pandas, NumPy
- Scikit-learn
- XGBoost
- Jupyter Notebook
- Hyperparameter tuning (GridSearch, Optuna)
- Time-series models (ARIMA, Prophet)
- Streamlit / FastAPI deployment
- Real-time prediction pipeline
A complete ML pipeline for sales forecasting with strong performance using Random Forest and feature optimization.