I've built production Generative AI systems at AWS and GPU-accelerated inference platforms at NVIDIA, improving model quality, latency, reliability, and cost across enterprise AI workloads.
At AWS, I productionized a RAG-based knowledge assistant using Amazon Bedrock, Knowledge Bases, Titan Embeddings, S3, and OpenSearch Serverless. By tuning chunking, metadata filtering, retrieval, and prompt context, I improved grounded-answer accuracy from 78% to 89% across 2,500 queries and reduced p95 retrieval latency by 30%.
I also built reusable foundation-model evaluation pipelines, reducing experiment turnaround by 60%, and introduced Guardrails, prompt management, regression testing, prompt caching, and intelligent prompt routing to reduce unsafe responses by 40%, regressions by 55%, and inference cost by 28%.
Before AWS, I worked at NVIDIA on CUDA/C++ and TensorRT inference systems, cutting end-to-end latency by 28%, improving throughput 2.3x through FP16/INT8 deployment, and delivering real-time perception workloads under 30 ms per frame. I've mentored 12 engineers on RAG architecture, prompt engineering, LLM evaluation, and production debugging.
