Pavan P
@pavanp
DevOps & SRE engineer with 4+ years automating AWS infrastructure, Kubernetes platforms, and mission-critical production reliability.
What I'm looking for
I’m a DevOps & Site Reliability Engineer with 4+ years of experience automating cloud infrastructure, building Kubernetes-based platforms, and operating mission-critical production systems at enterprise scale. I focus on reliability, automation, and platform engineering—especially where strong incident response and observability matter.
At Infosys on the American Express program, I supported the mission-critical American Express Identity and Apigee Platform (Tier 0) and helped maintain high availability across 40+ interdependent Java-based microservices in a 24×7 environment. I resolved production incidents, validated releases, executed production changes, and maintained SLA compliance, administering Apigee API Management and leveraging Git/GitHub for operational support.
As a Senior System Engineer, I built and optimized CI/CD pipelines with Jenkins and GitHub Actions, containerized microservices with Docker, and deployed workloads on Amazon EKS using Deployments, Services, Ingress, ConfigMaps, and Secrets via ArgoCD GitOps. I also engineered monitoring, logging, and alerting with Prometheus, Grafana, Node Exporter, Dynatrace, ELK, and CloudWatch, improving observability and reducing MTTR while handling SSL/TLS renewals and vulnerability remediation.
Most recently as a Technology Analyst (DevOps/SRE), I led P1/P2 Major Incident Management (MIM), driving rapid service restoration through RCA, disaster recovery (DR) exercises, capacity planning, and end-to-end observability. I provisioned AWS infrastructure with Terraform, implemented GitOps-based CI/CD with GitHub Actions and ArgoCD, managed CAB approvals and Emergency RFCs, and strengthened operational readiness through SOPs, post-incident documentation, and knowledge transfer. I was also recognized with the RISE INSTA Award (2×) for outstanding performance and operational excellence.
Experience
Work history, roles, and key accomplishments
Led P1/P2 major incident management and rapid service restoration using RCA, disaster recovery exercises, capacity planning, and production observability. Provisioned AWS infrastructure with Terraform and implemented GitOps-based CI/CD using GitHub Actions and ArgoCD, supporting CAB approvals, emergency RFCs, and production deployments.
Built and optimized CI/CD pipelines using Jenkins and GitHub Actions and deployed containerized microservices on Amazon EKS using Kubernetes primitives and ArgoCD GitOps. Engineered monitoring, logging, and alerting with Prometheus, Grafana, Node Exporter, Dynatrace, ELK, and CloudWatch to improve observability and reduce MTTR, and handled SSL/TLS certificate renewals and release validation.
Provided production support for the mission-critical American Express Identity and Apigee Platform (Tier 0), maintaining high availability across 40+ Java-based microservices in a 24x7 environment. Resolved production incidents, validated releases, executed production changes, and maintained SLA compliance while administering Apigee API Management and supporting operations via Git/GitHub.
Education
Degrees, certifications, and relevant coursework
Savitribai Phule Pune University
Bachelor of Technology, Computer Science and Engineering
2016 - 2020
Earned a B.Tech in Computer Science and Engineering from Savitribai Phule Pune University (Aug 2016–May 2020).
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring Pavan?
You can contact Pavan and 90k+ other talented remote workers on Himalayas.
Message PavanGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
