At Mphasis, I build end-to-end AWS data solutions, including a Synthetic Data Generation Automation Framework using Airflow, Lambda, Glue, S3, Redshift, and PySpark. I orchestrate metadata-driven Glue jobs and validate Parquet datasets ranging from 1M to 1B+ records.
I've developed PySpark ETL pipelines for hundreds of millions of financial records, loaded curated data into Snowflake for Power BI reporting, and reduced pipeline runtime by 40% through Spark optimization. I also automate CI/CD with Jenkins and Docker and implement monitoring, alerting, and data-quality checks.
