At Mckinsey & Company, I design and operate Spark and Python pipelines in Databricks, including production pipelines that process over 5 TB daily.
I’ve developed and maintained more than 15 SparkSQL jobs using medallion architectures and incremental data modeling. I also created dimensional models and schema designs that improved query performance by 20%.
To strengthen data quality and traceability, I implemented frameworks and lineage instrumentation using Monte Carlo, reducing data issues by 30%. I’ve also optimized Spark and SQL pipelines to reduce runtime by 30% and cloud compute costs by 25%.
At Progrite System, I designed reporting data models and supported production pipelines monitored with Airflow and Monte Carlo. I also supported Kafka and Kinesis streaming ingestion for event-driven reporting.

