At STE||AR-GROUP, I tuned and maintained the HPX backend of FleCSI as both projects evolved, and improved the performance of a distributed Poisson solver while ensuring correctness and scalability across runtime updates.
My work with HPX also includes adapting parallel algorithms to C++20, optimizing the rotate algorithm, and developing asynchronous models. At Lawrence Berkeley National Laboratory, I designed and implemented a GPU-accelerated Cholesky decomposition using standard C++ parallelism.

