Vrinda Marwah
@vrindamarwah
Software engineer focused on scalable Kubernetes and Slurm automation for AI/HPC platforms.
What I'm looking for
I’m a software engineer with 3+ years of experience in large-scale infrastructure automation and platform engineering. I specialize in designing, deploying, and managing highly available Kubernetes and Slurm clusters for AI/HPC workloads across air-gapped, multi-architecture environments (x86_64, ARM). I’m proficient in Ansible-driven automation and cluster lifecycle management, with a strong foundation in Linux and open-source development.
At Dell Technologies, I developed an Ansible-based infrastructure automation toolkit for HPC environments, deploying Slurm and Kubernetes clusters for multi-architecture Linux servers with network-based boot capabilities. I built an air-gapped local repository management system using Pulp to enable offline deployments while maintaining 99.9% package availability, and implemented diskless boot for Kubernetes and Slurm compute nodes to provision 100+ nodes simultaneously. I also engineered upgrade/rollback mechanisms with zero-downtime deployments and resolved complex cluster failures with root-cause analysis across OS, networking, and container runtime layers.
Previously, as an Intern at S&P Global, I deployed and managed Kubernetes clusters on AWS EKS, using blue-green deployment strategies and configuring Ingress and NodePort for service exposure. I built CI/CD pipelines to automate application deployment and scaling on AWS, reducing manual intervention and streamlining cloud operations workflows. I’ve also been recognized for innovation at the Dell ISG Hackathon, and I hold the NASSCOM Women Wizards Rule Tech (W²RT) certification, with a patent pending on context-aware AI/HPC orchestration.
Experience
Work history, roles, and key accomplishments
Developed an Ansible-based infrastructure automation toolkit for HPC environments, deploying Slurm and Kubernetes clusters for multi-architecture (x86_64, ARM64) air-gapped Linux servers. Built offline repository management and diskless boot capabilities, and implemented upgrade/rollback mechanisms and troubleshooting for cluster reliability.
Deployed and managed Kubernetes clusters on AWS EKS and supported blue-green deployments with load balancing and service exposure via Ingress and NodePort. Built CI/CD pipelines to automate application deployment and scaling in AWS.
Education
Degrees, certifications, and relevant coursework
University of Petroleum and Energy Studies (UPES)
B.Tech (Hons.), Computer Science and Engineering (Business Analytics and Optimization)
2019 - 2023
Grade: 9.25/10
B.Tech (Hons.) in Computer Science and Engineering with Business Analytics and Optimization, completed from 2019 to 2023 with a CGPA of 9.25/10.
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Job categories
Interested in hiring Vrinda?
You can contact Vrinda and 90k+ other talented remote workers on Himalayas.
Message VrindaGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
