At IBM, I was a key contributor on GEICO’s Policy Data Vertical, developing and maintaining Spark and Hive transformations that processed 500 GB to several TB per day. I converted legacy Hive and SQL transformations into Spark SQL and DataFrame jobs on YARN.
I diagnosed a skewed Spark join using the Spark UI and re-engineered it as a broadcast join, cutting runtime from about 100 to 65 minutes. I checked accuracy with row-count and aggregate reconciliation.
I owned the transformation layer of a roughly 20-table migration to Snowflake, building dbt staging and business models and Data Vault 2.0 mappings. At Capgemini, I built HiveQL transformations over raw HDFS data and performed Oracle SQL backend validation for GE Healthcare’s cardiology product.
I also built a serverless AWS streaming and batch data platform, and a RAG data pipeline with a data-quality gate before Qdrant indexing. I recently completed an MS in Artificial Intelligence at The University of Texas at Austin and hold AWS and Azure data engineering certifications.

