At Scale AI, I build the ingestion, validation, retrieval, and evaluation layers behind production LLM applications. I process roughly 40 TB per month of multimodal customer data and reduced re-labeling from bad source data by approximately 35%.
I’ve built embedding pipelines, pgvector and Milvus indexes, and hybrid RAG retrieval using BM25, dense search, reciprocal rank fusion, and cross-encoder reranking. This improved recall@10 from 0.71 to 0.89 while reducing ungrounded responses, and my telemetry-driven prompt caching, batching, and model routing helped reduce inference spend by approximately 30%.
Previously at DoorDash, I helped modernize real-time event ingestion with Kafka and Flink on Kubernetes, reducing analytics latency from roughly 24 hours to minutes across 200+ event tables. I also built systems handling billions of events per day at 99.99% delivery and mentored three junior engineers.
My earlier work at Optum and EPAM Systems grounded me in governed healthcare data, data quality, identity resolution, analytics, and business intelligence. I build production systems end to end with strong attention to reproducibility, observability, schema evolution, and secure data handling.
