Skip to main content
MIGHT UserMU
Open to opportunities

MIGHT User

@mightuser

Staff Site Reliability Engineer focused on reliable cloud platforms, Kubernetes automation, observability, and safe, fast deployments.

United States
Message

What I'm looking for

I’m looking for a team where I can build reliable, automated cloud platforms—improving SLO-driven observability, release safety, and incident recovery while reducing cloud costs and accelerating safe deployments.

I’m a Staff Site Reliability Engineer / Platform Engineer with 12+ years building and operating large-scale cloud infrastructure across TikTok, Snap, and Amazon. I specialize in Kubernetes reliability, infrastructure automation, and observability—always with the goal of improving uptime and speeding up safe deployments.

At TikTok, I built Flink-based aggregation and anomaly-detection pipelines and used Prometheus alerts to reduce regression detection time from ~20 minutes to under 5 minutes. I helped resolve ranking latency incidents with canary rollbacks and strengthened operational readiness through runbooks, RCA templates, and alert-routing guidance. I also improved service resilience with retries, timeouts, circuit-breakers, and fallbacks, and drove adoption of SLO dashboards and runbook patterns across sub-teams.

At Snap, I owned SLOs for Camera and AR platform services and built release-safety automation that combined CI/CD checks, synthetic validation, and staged rollout gates. I migrated AR backend services to ECS/EKS to support zero-downtime rollout practices, and I improved monitoring by creating Grafana dashboards and PagerDuty alerts to reduce detection time from hours to minutes.

Earlier in my career, I helped found Viro Media’s platform by implementing AWS backend infrastructure, Terraform-based self-service deployments, and scalable streaming services. My mindset stays the same: automate guardrails, measure outcomes with SLOs, and make incidents easier to detect, diagnose, and recover from.

Experience

Work history, roles, and key accomplishments

TikTok logoTI
Current

Staff Software Engineer

Dec 2023 - Present (2 years 7 months)

Built Flink-based aggregation and anomaly-detection pipelines with Prometheus alerts to monitor recommendation quality and reduce regression detection time from ~20 minutes to under 5 minutes. Improved release safety with canary rollbacks and runbooks, and reduced Kubernetes recommendation infrastructure costs via EKS right-sizing and autoscaling.

Snap logoSN

Senior Software Engineer

Nov 2019 - Dec 2023 (4 years 1 month)

Owned SLOs for Camera and AR platform services and built release-safety automation using CI/CD checks, synthetic validation, and staged rollout gates to catch regressions early. Migrated AR backend services to ECS/EKS and created Grafana dashboards and PagerDuty alerts to improve monitoring and incident detection.

Amazon logoAM

Software Development Engineer II

Jan 2015 - Aug 2016 (1 year 7 months)

Built synthetic canary checks and alerts for routing/geocoding vendors to monitor latency, error rate, data freshness, and route-quality regressions. Improved availability through multi-AZ migration components with circuit breakers and fallback logic, and automated seller-metrics data-quality remediation using event processing with DynamoDB and CloudWatch alarms.

Amazon logoAM

Software Development Engineer I

Aug 2013 - Jan 2015 (1 year 5 months)

Software Development Engineer I at Amazon Maps.

Education

Degrees, certifications, and relevant coursework

University of California, Berkeley logoUB

University of California, Berkeley

Bachelor’s Degree in Computer Science, Computer Science

2010 - 2013

Earned a Bachelor’s degree in Computer Science from the University of California, Berkeley from 2010 to 2013.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan