At 大众汽车安徽数字化销售有限公司, I manage production environments and service lifecycles for customer-facing microservices. I support releases, control deployment risk, and coordinate incident response.
I led upgrades to network topology and compute clusters, and planned highly available Kubernetes cluster upgrades. Across more than 100 business subsystems, self-service releases reached 95%+ and production release failures stayed below 1%.
I developed a Feishu-based operations troubleshooting platform that connects alerts, logs, release status, and AI-assisted analysis. It reduced the average time to investigate build failures and service incidents from about 30 minutes to under five.
I also established observability and reliability practices using Prometheus, VictoriaMetrics, Loki, and Grafana, and improved high availability for XXL-Job. I hold CKA, CKS, and SAFe 6 Agilist certifications.

