At DoorDash, I define evaluation and online measurement for LLM-powered discovery experiences, including conversational search and grounded store and item recommendations. I build evaluation datasets, compare LLM-as-judge scores with human review, and design experiments with GenAI-specific guardrails.
Previously at DoorDash, Best Buy, and eBay, I developed models and experiments for marketplace incentives, personalization, recommendations, and customer behavior. My work includes uplift modeling, causal evaluation, and measurement frameworks that connect product changes to business outcomes.

