At Alcon Laboratories, I built AWS Glue PySpark pipelines across HR and finance, processing 100M+ records daily. I redesigned deduplication logic to reduce downstream data volume by 26% and created a data health monitoring framework for production pipelines.
At E2open, I architected an AWS data lake pipeline using Airflow, Glue, S3, and Lambda, improving execution efficiency by about 30%. I also built a config-driven ingestion framework with schema validation and quarantine mechanisms to reduce data quality issues.

