In my self-directed data analysis practice, I transform claims-style feeds into canonical schemas using PySpark on Databricks, applying ICD-10, CPT and HCPCS coding. This work reduced downstream mapping errors by 45%.
I built Spark SQL and SQL validation checks that improved record-level quality to 99% across test datasets and cut manual validation time by 50%. I’ve also prototyped AI-assisted mapping workflows and claims payment integrity analyses.

