I developed processing logic for a Fraud Event Data Platform, working with Kafka-generated banking event data for Canada and the US. I transformed nested JSON into structured Delta datasets using reusable PySpark and Spark SQL transformations.
For the platform, I implemented incremental and backfill processing and added source-to-target reconciliation across data layers. I also troubleshot Delta, schema, and ADLS authentication issues and verified fixes with targeted reruns and regression checks.
In Enterprise ETL & Big Data, I developed ingestion workflows and created and maintained 50+ Hive tables, partitions, and HiveQL scripts. I also built RapidMiner mappings with Oracle sources and targets and performed SQL-based validation.
For Olympic Trends & Global Sports Analytics, I designed and built an Azure lakehouse pipeline using ADF, ADLS Gen2, Databricks, PySpark, Delta Lake, and Synapse. I connected curated datasets to Power BI for analysis of medal, country, sport, and gender trends.

