At BigBear.ai, I build and optimize distributed training and inference systems for generative AI, reducing end-to-end latency by 35%, improving throughput by 25%, and lowering GPU inference costs by 20% for production marketing workloads.
I design GPU-accelerated video and image generation pipelines, agentic AI systems for customer data enrichment, low-latency backend APIs, and scalable MLOps workflows. My work includes model compression, quantization, benchmarking, profiling, explainability, monitoring, and multi-GPU optimization for reliable customer-facing AI products.
Previously, I delivered machine learning and forecasting solutions at Johnson & Johnson, productionized ranking and recommendation models at Facebook, and built high-performance database systems at Comcast. I also mentor engineers on GPU optimization, distributed systems, and experimental AI research.
