At Amazon, I build end-to-end batch and event-driven data platforms for mission-critical datasets. I delivered pipelines using AWS Glue, Apache Airflow, and PySpark that achieved 99.9% data accuracy and improved processing performance by 40%.
I design canonical datasets, data contracts, quality guardrails, and BI foundations for reporting, experimentation, and real-time use cases. I also led Redshift deprecation across 15+ teams, 300+ customers, and 600+ pipelines, driving roughly 40% infrastructure cost savings while mentoring engineers and analysts.

