This project implements a fully automated, containerized, and scheduled data engineering pipeline to perform RFM (Recency, Frequency, Monetary) analysis on e-commerce transaction data. The system automatically ingests raw data from Google BigQuery, processes it through a complex SQL transformation, and saves the final customer segments back to a BigQuery table on a daily schedule.
This project demonstrates the transformation of a manual data analysis task into a robust, production-ready, and hands-off engineering solution.
This project is the automated, production-ready version of an initial exploratory analysis. The original work, which involved manual, standalone SQL queries to derive business insights, can be found in the RFM Customer Segmentation & Retention Analysis.
This pipeline takes the core logic from that analysis and re-engineers it into a robust, scheduled, and containerized system suitable for a production environment.
- Cloud: Google Cloud Platform (GCP)
- Data Warehouse: Google BigQuery
- Orchestration: Python
- Containerization: Docker
- Scheduling: Windows Task Scheduler / cron
- Language: SQL
The pipeline follows a simple, automated workflow:
- Scheduled Trigger: A system scheduler (like cron or Windows Task Scheduler) kicks off the process daily.
- Docker Execution: The scheduler runs a
docker runcommand to start the containerized application. - Python Orchestration: The
pipeline.pyscript takes over, managing the connection to BigQuery and the execution of the SQL logic. - SQL Transformation: A single, efficient SQL query with CTEs performs all the data cleaning, RFM calculation, scoring, and segmentation in one pass.
- Load to BigQuery: The final results are saved to a destination table in BigQuery, ready for use by business intelligence tools or marketing teams.
- Docker Desktop installed and running.
- A Google Cloud Platform account with a project set up.
- A
credentials.jsonfile from a GCP Service Account with BigQuery permissions.
- Clone this repository.
- Place your
credentials.jsonfile in the root of the project directory. (Note: This file is included in.gitignoreand will not be committed). - Update the
PROJECT_IDandDESTINATION_TABLEvariables inpipeline.pywith your GCP project details.
-
Build the Docker Image:
docker build -t rfm-pipeline . -
Run the Container:
docker run --rm rfm-pipeline
The pipeline will execute and the final results will be available in your specified BigQuery destination table.