Skip to main content
Muhammad EhtishamME
Looking for a job

Muhammad Ehtisham

@ehtishammubarik

Senior AI platform engineer building agentic LLM systems and scalable cloud/Kubernetes infrastructure with cost and reliability focus.

United States
Message

What I'm looking for

I’m looking to lead production AI platform work—agentic LLM systems and distributed ML on Kubernetes—while owning reliability, security/compliance, and measurable cost optimization. I want teams that ship fast, test hard, and treat SRE as a first-class capability.

I’m a Senior AI Platform Engineer with 7+ years building and shipping production AI and cloud infrastructure. I focus on agentic systems that actually run in production—subagent orchestration, dynamic workflows, custom skill packaging, and token-cost optimization for LLM and agent workloads.

In my recent work, I rebuilt production AWS data and ML around a Kafka-driven pub/sub architecture for a US gaming platform—Bedrock, Airflow, Redshift, and Glue. I delivered 60%+ AWS spend reduction through smarter model selection and inference harness improvements, disciplined rightsizing, and senior-level operational best practices. I also engineered reliable and secure production layers with layered IAM/RBAC, secrets management, image hardening, network policy enforcement, and ongoing regression/load/stress plus periodic penetration testing.

I’ve expanded that approach across multi-cloud and on-prem GPU orchestration on Kubernetes, using Ray-based distributed training/inference and Kubernetes automation tooling. From Stack8s to consulting engagements, I’ve designed multi-tenant cluster provisioning (GPU pools, Gateway API routing, Rancher management) and built Pulumi/Go modules to automate per-tenant infrastructure and teardown. I bring production SRE leadership—PagerDuty escalation, SLA/SLO enforcement, incident triage harnesses—and I’m CKA and CKAD certified.

Experience

Work history, roles, and key accomplishments

HL
Current

Senior DevOps Engineer

Hyve Labs

Apr 2025 - Present (1 year 3 months)

Owned end-to-end production management for a gaming platform, including real-time event fan-out and elastic multi-region autoscaling. Rebuilt the AWS data and ML stack around Kafka pub/sub, reduced AWS spend by 60%+ on Bedrock, and implemented secure production reliability and incident operations.

ST
Current

Senior DevOps Engineer

Stack8s

Sep 2024 - Present (1 year 10 months)

Worked on the Stack8s Kubernetes automation platform for AI and ML infrastructure teams. Built multi-cloud/on-prem GPU-enabled cluster provisioning with Kubeflow, Kamaji-based multi-tenancy with Gateway API routing, and Rancher-based cluster management.

DC

Sr DevOps & MLOps Consultant

Dressler Consulting

Nov 2022 - Jul 2026 (3 years 8 months)

Designed Ray-based distributed architectures for AI training and inference across on-prem and cloud-managed environments. Built Pulumi (Go) modules for a multi-tenant AI platform, owned alerting for distributed training, and developed an agentic harness for orchestration and token-optimized long-running jobs.

AA

DevOps Engineer, Team Lead

Aceso Analytics

Apr 2020 - Sep 2024 (4 years 5 months)

Led an engineering team on a healthcare IoT platform serving LoRaWAN device fleets across elderly care facilities. Architected the IoT and LoRaWAN stack on Kubernetes with HIPAA-aligned operational practices and built scalable monitoring, alerting, and edge inference, mentoring junior engineers.

AF

Software Engineer DevOps

Afiniti

Oct 2020 - Dec 2021 (1 year 2 months)

Automated build, packaging, and release of 40+ microservices across air-gapped enterprise environments. Migrated legacy infrastructure to on-prem Kubernetes using production Ansible playbooks and Jenkins shared libraries.

HS

DevOps Engineer

Halfpenny Software

Jan 2019 - Dec 2021 (2 years 11 months)

Provided parallel DevOps engagements across US and international clients on AWS, Azure, and GCP. Led containerization and clustering migrations using Terraform, Ansible, and Bash.

EM

DevOps Engineer

Emumba

Jul 2019 - Oct 2020 (1 year 3 months)

Built monitoring dashboards and observability stacks for on-prem and cloud customers. Developed cross-platform ticket routing and custom CI/CD runners to support customer delivery workflows.

Education

Degrees, certifications, and relevant coursework

National University of Sciences and Technology (NUST) logoNN

National University of Sciences and Technology (NUST)

Master of Science in Computer Science, Computer Science

2019 - 2021

Master of Science in Computer Science at NUST (2019–2021). Thesis: GPU Cluster Optimization for ML Workloads.

National University of Computer and Emerging Sciences (FAST-NU) logoNF

National University of Computer and Emerging Sciences (FAST-NU)

Bachelor of Science in Computer Science, Computer Science

2015 - 2019

Activities and societies: Organized internal AI/MLOps/agentic-tooling workshops, attended cloud & infrastructure meetups, and published on production AI/agentic system operations.

Bachelor of Science in Computer Science at FAST-NU (2015–2019). Participated in organizing internal AI/MLOps/agentic-tooling workshops, attending cloud/infrastructure meetups, and publishing on production AI and agentic system operations.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan