At Baseten Labs, Inc., I operate production GPU inference infrastructure across 70+ Kubernetes clusters and ~20 cloud providers. I built production Kubernetes clusters and the GitOps workflows that provision networking, GPU enablement, and autoscaling for customer AI workloads.
I also built GPU hardware triage playbooks that cut time-to-disposition for failing nodes from days to hours. I carry P0 on-call for the serving fleet, leading incident command and root-cause isolation through restoration.
At Observe.AI, I founded the SRE platform and built centralized observability and alerting across AWS and Kubernetes. I also drove cloud right-sizing that resulted in approximately $15,000 per month in sustained savings, and designed AWS SageMaker infrastructure for production ML workloads.

