You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Star1 (1)You must be signed in to star a repository
About
A production-grade, distributed ETL pipeline for processing application log data at scale using PySpark and a Medallion (Bronze/Silver/Gold) architecture — compatible with Databricks, AWS EMR, Azure Synapse, and local Spark clusters.
A production-grade, distributed ETL pipeline for processing application log data at scale using PySpark and a Medallion (Bronze/Silver/Gold) architecture — compatible with Databricks, AWS EMR, Azure Synapse, and local Spark clusters.
# In a Databricks notebook:%run/path/to/src/pipelinerun_pipeline(
input_path="dbfs:/mnt/raw/logs/",
output_path="dbfs:/mnt/processed/",
env="databricks"
)
Parquet — columnar output format with partition pruning
About
A production-grade, distributed ETL pipeline for processing application log data at scale using PySpark and a Medallion (Bronze/Silver/Gold) architecture — compatible with Databricks, AWS EMR, Azure Synapse, and local Spark clusters.