At C-DAC, I debug and analyze GPU-to-GPU communication and GPU-aware MPI code paths, investigating data-transfer inefficiencies and topology-aware strategies.
I built a simplified MPI library from scratch in C++ using TCP/IP sockets, and developed a multithreaded task scheduler with synchronization primitives. I also parallelized a bio-inspired optimization algorithm with OpenMP and CUDA, achieving up to 10x improvement in solution accuracy over baseline methods.

