At Optum, I led the design and implementation of a cloud-native Lakehouse platform that brings clinical, provider, claims, and operational data into a governed enterprise analytics platform. I also modernized legacy ETL processes into reusable metadata-driven ELT frameworks, reducing maintenance effort by 45%.
I designed Spark and PySpark pipelines for healthcare records and built batch and real-time streaming solutions for operational reporting. My work also includes developing AI-ready datasets and feature engineering pipelines for machine learning, LLM, and RAG initiatives.
At Capital One, I designed data platforms supporting fraud detection, regulatory reporting, customer analytics, and financial risk management. I built cloud-native pipelines and streaming architectures, and helped migrate legacy workflows to Spark-based processing.
Earlier at IBM, I developed enterprise ETL pipelines and reusable Python frameworks for data ingestion. I also built Spark transformations and automated workflow orchestration with Apache Airflow.

