At Tesla, I led training and productionization for an end-to-end multimodal foundation-model workstream during the Cortex/V13 scale-up. I re-architected distributed training so multi-day runs could recover after worker or GPU failures instead of restarting from a full snapshot.
I also turned bespoke training scripts into reproducible Docker and Kubernetes launches with Argo Workflows and MLflow lineage. My release path connected offline regression, shadow validation, and staged rollout so deployments could be reproduced and rolled back to an exact known state.
At Tesla, I built a fleet-scale hard-example mining and auto-labeling system for FSD perception. It used embeddings, FAISS/ANN search, and uncertainty-based sampling to find rare production failures, then routed ambiguous sequences to human reviewers and high-confidence examples into training.
At Meta, I helped build Looper, a self-service real-time ML optimization platform, and built retrieval and ranking components for Instagram Suggested Posts. Earlier, I embedded a lightweight C++ classifier in Facebook's reactive-cache path, replacing manually maintained filtering rules with a continuously adapting prediction and retraining loop.

