At the State of Ohio, I design SparkSQL pipelines in Python and maintain Apache Spark workflows that process 10 million+ records daily for operational and risk analytics.
I developed dimensional data models and medallion architectures, creating 15+ fact and dimension tables for datamarts. I also implemented Databricks ETL workflows with a cloud data lake, reducing refresh latency by 30% for compliance-driven reporting.
I orchestrated Airflow DAGs for data quality monitoring using Monte Carlo methodologies, reducing data errors by 40%. I optimized SQL and SparkSQL queries and enhanced lineage instrumentation and testing frameworks.
Previously, at Sagarsoft Inc. for client JPMorgan Chase & Co., I developed Spark and Python pipelines on Databricks and modeled reporting marts. Earlier, at Aditya Birla Capital Limited, I worked on SparkSQL pipelines, cloud data lake migrations, and Kafka and Kinesis streaming data for operational dashboards.

