At Persistent Systems Limited, I migrated healthcare data schemas from PCORnet CDM v6.1 to v7.0 across five partner sites. I also led the migration from Hadoop MapReduce to Apache Spark, reducing average job processing time by 50% and cluster compute costs by 30%.
Earlier, I developed ETL workflows and data warehouse solutions for healthcare data, including applying Datavant hashing and tokenization to patient records for HIPAA compliance. I also developed serverless APIs and automated data processing with AWS services, and explored genomic sequencing datasets to improve data quality metrics.

