At Microsoft, I lead a distributed SRE organization responsible for reliability, escalation response, and operational excellence across hybrid cloud and on-premises data platforms. I own high-severity incident response from triage through post-incident reviews and structural fixes.
I built monitoring and observability programs for enterprise customers using SQL/KQL analytics and Power BI/Tableau reporting. I sustained 99.95%-99.98% actual platform reliability across 14+ consecutive weeks and improved perceived reliability from 98.63% to 99.67%-99.82%.
I also designed and delivered AI-assisted triage workflows for incident analysis (Cortex/MCP), with 80% accuracy and high confidence scores, now used by partner teams. Earlier, I delivered cloud migration and platform engineering solutions for HDInsight workloads, and contributed code-level fixes to Apache Hive at Cloudera.

