At Spring Health, I architected clinical knowledge retrieval with production RAG pipelines on FAISS and Pinecone, improving retrieval accuracy by 35% while serving queries at sub-200ms p95 latency.
I designed multi-agent LLM workflows for mental health intake that reduced clinician intervention by 40% and administrative workload by 25%. I also fine-tuned Llama 3 and Mistral with LoRA and QLoRA on de-identified clinical datasets, gaining 28% in contextual accuracy.
At Google, I built and deployed TensorFlow models and production ML pipelines powering real-time analytics across internal products. I also designed NL-to-SQL pipelines and helped bring pretrained transformer models and vector search into use across partner teams.
