At Porter, I build PySpark pipelines on AWS EMR that turn data from Amazon S3 into analytics-ready datasets. I also organize data into raw, cleansed, and curated zones.
I develop reusable transformations and full-load or incremental-load workflows based on business requirements. I separate rejected records from valid data and store validation failures in Amazon S3.
I design Apache Airflow DAGs and use AWS Step Functions to orchestrate Spark workloads. I also configure EMR clusters, optimize Spark jobs, and debug processing errors.

