At Activeloop, I own the architecture, design, and implementation of Deep Lake's high-performance C++ data-loading subsystem. I delivered more than 10x improvement in data-loading throughput while maintaining efficient memory use and maximizing GPU utilization during AI training workloads.
I also designed multithreaded and multiprocess data pipelines, cloud-native dataset access, and storage-layer optimizations. I integrated Deep Lake with MMDetection and MMSegmentation, and developed backend C++ infrastructure and Python-facing APIs used by thousands of users.
At Xilinx, I contributed to Vivado FPGA routing technologies and performance-critical C++ infrastructure for routing scalability and quality. I also participated in algorithmic improvements for NP-hard routing optimization problems.
At Softhenge, I led the redevelopment of Energi's GPU mining infrastructure using CUDA and OpenCL, and designed public mining pool infrastructure. Earlier, I contributed to EDA tools at Mentor Graphics (Siemens EDA) and developed C++/Qt interface components at ScorpTech.

