At ZS Associates, I build production-grade AWS data platforms for pharmaceutical commercial and clinical data. I've engineered ETL/ELT pipelines processing 4TB+ per run and unified multi-vendor Marketing and Sales datasets into a data lake and warehouse.
I've improved Apache Spark processing performance by 60%, saving 4,000+ compute hours annually, through broadcast joins, predicate pushdown, partition pruning, and parallel computing. I also reduced pipeline runtime by approximately 45% through OLAP-optimized data modeling and query tuning.
I modernize data operations through Kedro, Argo Workflows, Kubernetes, CI/CD, automated testing, governance, reconciliation, and lineage tracking. My work has supported oncology targeting for 50,000+ healthcare providers while maintaining audit-ready, reliable data products.
