At NVIDIA, I build distributed LLM training, evaluation, and inference capabilities on DGX Cloud, optimizing transformer workloads, recovery mechanisms, model-regression gates, and GPU-aware serving.
Previously at Tesla, I developed machine-learning systems for manufacturing quality at Giga Texas, turning equipment, process, inspection, and component-genealogy data into anomaly and quality-risk predictions integrated with auditable engineering workflows.
At Google, I built diagnosis and recommendation capabilities for Cloud Network Intelligence Center using dependency graphs, classification, ranking, and evidence-grounded explanations.
At Arista Networks, I developed telemetry analytics and anomaly detection for CloudVision, helping operators investigate network events through prioritized anomalies and time-series evidence.
I work across the full AI lifecycle: data preparation, feature engineering, distributed training, evaluation, deployment, monitoring, and continuous improvement. I focus on translating advanced AI research into reproducible, scalable production software.
