At Qorvo, I optimize production LLM systems across multi-GPU Linux environments, model serving APIs, and performance monitoring. I reduced per-epoch training time by 35%, production inference latency by 40%, and GPU memory utilization by 25%.
I build reproducible ML services with PyTorch, FastAPI, Docker, MLflow, GitHub Actions, SQL, Parquet, and FAISS. My work spans distributed training, quantization, KV-cache reuse, request batching, streaming, caching, benchmarking, and production validation.
I also built an LLM Inference Optimization Benchmark Suite that achieved up to 3x throughput improvement and 60% GPU memory reduction, plus a multi-agent SQL analyst that turns business questions into validated SQL and reports. Earlier at DXC Technologies, I developed Python and SQL data pipelines, feature-engineering workflows, supervised-learning models, and automated data-quality monitoring.

