At KwikQuery LLC, I started the TabbyDB venture and developed an Apache Spark fork that addresses complex query compilation issues and improves runtime performance with dynamic file pruning for Broadcast Hash Joins. Internal TPC-DS testing on 1 TB and 3 TB datasets showed 35% improvement compared with stock Spark.
At Workday Inc, I implemented a constraint propagation algorithm for Spark’s Catalyst optimizer that reduced compilation time by a factor of 10–100 in problematic cases. I also proposed an improvement to open-source Spark and presented the solution at Databricks Spark Summit 2021.
Across my work at Cloudera Inc, SnappyData, Pivotal Labs, and GemStone Systems, I contributed to Spark SQL and Catalyst, distributed databases, and query engines. My projects included approximate query processing, Spark and GemFireXD integration, and Java systems using concurrency and low-latency techniques.

