SparkPySparkETLAWS
Big Data Analytics Pipeline
Distributed ETL and analytics on large-scale datasets using Spark and cloud storage.

A distributed data pipeline that ingests, transforms, and serves analytics on large-scale datasets. Built with Apache Spark and deployed on cloud storage and compute.
Scope
The pipeline processes terabytes of event data daily. We use PySpark for transformations, partition data by date and key dimensions, and expose aggregated metrics for dashboards and APIs.
Tech stack
Apache Spark for batch and micro-batch jobs, S3-compatible storage for raw and curated layers, and Airflow for orchestration. All jobs are idempotent and support incremental processing.
Gallery

