At Granica, I built a columnar storage engine for Apache Parquet, with pluggable readers, encoding selection, and ACID commit paths for Hive, Delta Lake, and Iceberg. It cuts storage 20–40% with byte-exact output.
I also scaled distributed ingestion from 1.06 to 15.6 GB/s and redesigned a commit protocol so commits went from about 30 minutes to seconds. I built a Python and FastAPI control plane, Kubernetes orchestration, and a test platform that certifies releases against live AWS and GCP environments.
Previously, at Rippling, I led the migration of the payroll amendment pipeline to new infrastructure and developed deterministic lifecycle semantics and audit tooling. At Cloudera, I optimized distributed query engines and storage formats, contributed upstream to Apache Hive, Tez, and Hadoop, and built tooling for Hive Metastore and runtime log analysis.

