At Innovaccer, I designed and maintained end-to-end ETL/ELT pipelines for healthcare analytics and reporting using Python, SQL, and PySpark. I automated ingestion and processing of 1TB+ of data daily, reducing manual effort by 90% and increasing pipeline throughput by 50%.
I optimized ETL workflows through query tuning and efficient cloud resource utilization, improving processing speed by 60% and reducing infrastructure costs by approximately $10K. I also developed Python workflows for file extraction and retrieval from Amazon S3.
I contributed to data migration across Elasticsearch, PostgreSQL, and Snowflake using Apache Airflow, and built data validation frameworks that reduced processing errors by 90%. I developed and executed 100+ SQL validation scripts and ETL test cases, and implemented monitoring and error handling for ingestion pipelines.
At PrepInsta, I processed and analyzed structured datasets using Python and Pandas and developed Power BI dashboards. In my Netflix Data Pipeline project, I built an ELT pipeline using Amazon S3 and Snowflake, with a multi-layer warehouse and five analytics-ready mart tables.

