At Bigeye, I architect and scale multi-cloud data platforms across AWS, Azure, and GCP. The platforms integrate batch and streaming pipelines and improved analytics latency by 45%.
I lead cross-functional engineering and data science teams, and introduced governance, data quality, and cost-optimized cloud infrastructure. With Collibra, Atlan, and Monte Carlo, I implemented lineage, cataloging, and observability that achieved 99.9% data reliability.
At Starburst, I designed ETL/ELT pipelines for healthcare and clinical datasets, reducing processing latency by 50%. I also built Kafka and Spark Streaming pipelines for near-real-time monitoring of hospital operations and patient outcomes.
Earlier at Cloudera, I built and optimized data pipelines and streaming systems, and developed data models and dimensional schemas. My work included improving pipeline runtime and supporting Hadoop and Hive workflow migrations to AWS and GCP.
