At Accenture, I cut a 5TB fact-table load from 60 to 35 minutes and improved BigQuery/Dataflow processing time by 20% while reducing operational cost by 15%.
I've built production batch and streaming ETL with PySpark, Spark, Databricks, BigQuery, Apache Beam, Dataflow, Dataproc, Airflow, and Lakeflow. I also developed an AI metadata agent that reduced manual source-to-target mapping effort from roughly two days to 25 minutes and generated reference Lakeflow pipelines through MCP.
At Flipkart, I engineered BigQuery ETL for datasets serving a 500M+ user base, improving load efficiency by 35%. I delivered ML feature pipelines for production models and used PySpark streaming for peak-traffic data, reducing latency and improving downstream model true-positive rate from 30% to 38%.
Earlier, I established Interview Kickstart's data and analytics function, delivering pipelines, governance, quality frameworks, and real-time dashboards. I'm a Databricks Certified Data Engineer Professional and Google Professional Data Engineer focused on reliable, well-modeled data that teams can trust.
