MIGHT User
@mightuser
Staff Site Reliability Engineer focused on reliable cloud platforms, Kubernetes automation, observability, and safe, fast deployments.
What I'm looking for
I’m a Staff Site Reliability Engineer / Platform Engineer with 12+ years building and operating large-scale cloud infrastructure across TikTok, Snap, and Amazon. I specialize in Kubernetes reliability, infrastructure automation, and observability—always with the goal of improving uptime and speeding up safe deployments.
At TikTok, I built Flink-based aggregation and anomaly-detection pipelines and used Prometheus alerts to reduce regression detection time from ~20 minutes to under 5 minutes. I helped resolve ranking latency incidents with canary rollbacks and strengthened operational readiness through runbooks, RCA templates, and alert-routing guidance. I also improved service resilience with retries, timeouts, circuit-breakers, and fallbacks, and drove adoption of SLO dashboards and runbook patterns across sub-teams.
At Snap, I owned SLOs for Camera and AR platform services and built release-safety automation that combined CI/CD checks, synthetic validation, and staged rollout gates. I migrated AR backend services to ECS/EKS to support zero-downtime rollout practices, and I improved monitoring by creating Grafana dashboards and PagerDuty alerts to reduce detection time from hours to minutes.
Earlier in my career, I helped found Viro Media’s platform by implementing AWS backend infrastructure, Terraform-based self-service deployments, and scalable streaming services. My mindset stays the same: automate guardrails, measure outcomes with SLOs, and make incidents easier to detect, diagnose, and recover from.
Experience
Work history, roles, and key accomplishments
Built Flink-based aggregation and anomaly-detection pipelines with Prometheus alerts to monitor recommendation quality and reduce regression detection time from ~20 minutes to under 5 minutes. Improved release safety with canary rollbacks and runbooks, and reduced Kubernetes recommendation infrastructure costs via EKS right-sizing and autoscaling.
Owned SLOs for Camera and AR platform services and built release-safety automation using CI/CD checks, synthetic validation, and staged rollout gates to catch regressions early. Migrated AR backend services to ECS/EKS and created Grafana dashboards and PagerDuty alerts to improve monitoring and incident detection.
Founding Engineer
Viro Media
Aug 2016 - Aug 2019 (3 years)
Built core AWS backend infrastructure using multi-AZ ECS, CloudFront, and DynamoDB-backed services to support large-scale AR traffic. Implemented Terraform modules and self-service deployment pipelines, added CloudWatch monitoring/runbooks, and improved platform security with rate limiting and AWS Shield/WAF rules.
Built synthetic canary checks and alerts for routing/geocoding vendors to monitor latency, error rate, data freshness, and route-quality regressions. Improved availability through multi-AZ migration components with circuit breakers and fallback logic, and automated seller-metrics data-quality remediation using event processing with DynamoDB and CloudWatch alarms.
Software Development Engineer I at Amazon Maps.
Education
Degrees, certifications, and relevant coursework
University of California, Berkeley
Bachelor’s Degree in Computer Science, Computer Science
2010 - 2013
Earned a Bachelor’s degree in Computer Science from the University of California, Berkeley from 2010 to 2013.
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring MIGHT?
You can contact MIGHT and 90k+ other talented remote workers on Himalayas.
Message MIGHTGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
