At Google, I lead GKE AI infrastructure across three product surfaces and guide a team of six engineers. I was the founding engineer for TPU support on Google Kubernetes Engine.
I shipped managed TPU v4, v5e, v5p, and Trillium support across tens of thousands of schedulable chips running JAX, XLA, and PyTorch/XLA workloads. I redesigned the Go node-pool lifecycle path, reducing provisioning from 18 to 6 minutes and time to first training job by 65%.
Previously, I built GKE control-plane services for a six-figure cluster fleet and reduced cluster-upgrade failures by 40%. At Amazon, I worked on exabyte-scale S3 durability and control planes, Alexa self-learning infrastructure, and Aurora database internals.
My work spans Go, Kubernetes, gRPC, cloud control planes, storage systems, and AI infrastructure. I’ve built systems that move petabytes per day, serve tens of millions of daily requests, and support production machine-learning workloads at scale.

