At IBM’s watsonx.data Spark SaaS team, I designed and implemented an event-driven teardown mechanism that eliminated about $127K per month in orphaned Spark infrastructure costs.
I also engineered automatic retries for transient Spark failures, reducing recovery time to about two minutes across more than 5,000 daily workloads. By resolving API reliability and performance issues, I reduced production errors by 89.4% and p99 latency by 19.2%.
I architected REST APIs for Spark compute capacity management and built secure private access to Spark runtimes, helping migrate more than 25 enterprise customers. I also took a proof of concept for generic S3-compatible storage registration through to production.
I contribute to Apache Spark, including changes to join pushdown, SQL parser diagnostics, and the Spark Web UI. On my Helix project, I built a Java JVM diagnostic platform and an LLM agent extended through an MCP server.

