At xAI, I build Grok’s reasoning and tool-use capabilities through RL post-training, reward modeling, and agentic curricula. I engineered visual-grounding verifiers for Grok 4.1 that cut hallucination rates by approximately 2x on production information-seeking prompts.
Previously at Google DeepMind, I built multimodal long-context evaluations and tool-use harnesses for Gemini 1.5 through 2.5, including thinking-model evaluation, code-execution sandboxes, and controllable thinking-budget calibration. At Google Research, I developed video-language systems behind MTV, VideoCoCa, and VideoPrism using JAX, Flax, Scenic, TPU infrastructure, and rigorous video evaluation.
At Meta AI, I owned PyTorch and VISSL pipelines supporting SEER, DINO, PAWS, MAE, and TimeSformer. I work across model training, evaluation, distributed systems, and production-quality failure analysis to make advanced multimodal models more reliable.

