At GoodRx, I architect Spark Structured Streaming pipelines with Kafka ingestion that process over 1 billion daily events and make data available in under a minute. I also orchestrate Airflow workflows integrating Databricks and Spark SQL tasks.
At ConstructionBevy, I developed Python ingestion pipelines and SQL schemas that reduced data onboarding time by 60%. I also automated data quality monitoring and implemented incremental data modeling and medallion architecture patterns.
In my projects, I designed an Apache Beam pipeline for London Cycles data and built a PySpark recommendation engine on AWS EMR. I also developed a real-time market data pipeline using Kafka and Spark Structured Streaming.

