At GAIL, I deployed and operated an on-premise speech-to-text inference stack on Kubernetes serving 3M+ calls per day. I optimized GPU placement, autoscaling, and network locality, and built a self-hosted Qwen3-8B inference service that reduced external API costs by 70%.
At Cornell BURE Summer Research Program, I built a distributed ML experimentation pipeline over 400k+ samples using Spark on a SLURM-managed HPC cluster. My experiments included Transformer, MLP, and gradient-boosting pipelines, with the best classifier achieving 92% accuracy at a 40-citation threshold.
With JULI, an AI app store for agents, I built an orchestration layer for multi-step workflows and deployed the backend on AWS. I also reimplemented Cold Diffusion in PyTorch for image restoration, improving RMSE by 4.4%; the project was voted Best Project by the course TAs.

