Skip to main content
sheshank dudaboinaSD
Looking for a job

sheshank dudaboina

@sheshankdudaboina

Site Reliability Engineer at Baseten operating GPU inference across 70+ Kubernetes clusters and ~20 cloud providers.

United States
Message

At Baseten Labs, Inc., I operate production GPU inference infrastructure across 70+ Kubernetes clusters and ~20 cloud providers. I built production Kubernetes clusters and the GitOps workflows that provision networking, GPU enablement, and autoscaling for customer AI workloads.

I also built GPU hardware triage playbooks that cut time-to-disposition for failing nodes from days to hours. I carry P0 on-call for the serving fleet, leading incident command and root-cause isolation through restoration.

At Observe.AI, I founded the SRE platform and built centralized observability and alerting across AWS and Kubernetes. I also drove cloud right-sizing that resulted in approximately $15,000 per month in sustained savings, and designed AWS SageMaker infrastructure for production ML workloads.

Experience

Work history, roles, and key accomplishments

Baseten Labs, Inc. logoBI
Current

Site Reliability Engineer

May 2026 - Present (5 months)

Operate production GPU inference infrastructure across 70+ Kubernetes clusters on ~20 cloud providers, spanning thousands of NVIDIA H100, B200, and A100 nodes. Built production Kubernetes clusters for customer AI workloads through a GitOps workflow and deployed the in-cluster platform stack for model serving and networking.

Observe.AI logoOB

Lead Site Reliability Engineer

Aug 2024 - May 2026 (1 year 9 months)

Built and scaled the company's first platform-level SRE function, owning reliability, observability, cost efficiency, and production operations across multi-region AWS and Kubernetes environments. Founded the SRE platform from first principles, defined on-call rotations, incident response workflows, and alerting standards.

Apple logoAP

AWS Site Reliability Engineer

Jun 2024 - Aug 2024 (2 months)

Contributed to the planning phase of an infrastructure migration project for payment platform applications, transitioning on-premises workloads from a data center to AWS. Collaborated with cross-functional teams to design a scalable, secure, and highly available architecture for payment applications on AWS.

Venmo logoVE

Site Reliability Engineer

Venmo

Oct 2018 - May 2019 (7 months)

Designed and operated AWS infrastructure at scale using Terraform, including networking, IAM, and multi-region deployments, for high-throughput financial transaction platforms. Built Kubernetes platforms (KOPS) and integrated them with GitLab CI/CD to support containerized microservices.

Education

Degrees, certifications, and relevant coursework

New England College logoNC

New England College

MBA, Business Administration

2022 - 2023

MBA coursework at New England College, not completed.

New England College logoNC

New England College

Master of Science, Information Technology & Cybersecurity

Master of Science in Information Technology & Cybersecurity from New England College, completed in 2022.

Stratford University logoSU

Stratford University

Master of Science, Information Science

2016 - 2017

Master of Science in Information Science from Stratford University, completed in 2017.

Kakatiya University logoKU

Kakatiya University

Bachelor of Technology, Information Technology

2011 - 2015

Bachelor of Technology in Information Technology from Kakatiya University, completed in 2015.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan