At Cognizant, I build and operate AWS data pipelines for Novartis, processing 1M+ healthcare records daily across 20+ ETL workflows. I've optimized PySpark transformations to reduce ETL runtime by up to 40% and improved Redshift performance by 20–25%.
I design reliable batch and streaming data platforms using Python, PySpark, AWS Glue, S3, Redshift, Kinesis, and Airflow. I also built a near-real-time global air-quality analytics platform with a governed Bronze/Silver/Gold data lake, dimensional modeling, and automated workflow recovery.

