At Baker Hughes, I build CDC-driven ingestion pipelines with Azure Data Factory, PySpark, and Databricks, bringing databases, APIs, and files into Unity Catalog and ADLS.
I've created reusable SQL ingestion and query-planning frameworks for Silver and Gold data layers, accelerating pipeline development by 40% while improving ETL efficiency. I also optimized Spark workloads to deliver 45% faster runtimes and 30% cost savings.
Previously at Wipro Technologies, I built and optimized Azure-based ETL pipelines for US banking and financial services clients, reducing processing time by 25% and improving reporting performance by 20% through scalable data models.
I enjoy making data platforms more reliable and governed, from automated validation and anomaly detection to near-real-time HVR replication. My projects include a PowerApps-to-Databricks connector, an in-house PySpark ETL framework, and a metadata lineage and testing portal integrated with Power BI.

