At Scale AI, I own end-to-end development of an enterprise LLM evaluation platform, increasing benchmark coverage by 42% and reducing regression detection time by 35%.
I build automated and human-in-the-loop evaluation workflows with Python and PyTorch, including RAG evaluation, agentic workflow testing, AI safety benchmarks, and LLM-as-Judge frameworks.
Previously at NVIDIA, I deployed GPU-accelerated LLM inference microservices with Kubernetes, Docker, and TensorRT-LLM, improving throughput by 45% and reducing production latency by 32%.
I also bring production ML, cloud, and data-platform experience from Accenture, where I built AI-powered applications, data pipelines, REST APIs, and cloud-native deployments across AWS and Azure.

