I've evaluated conversational AI agents in the confidential Lyra Project, testing complex multi-turn interactions and designing adversarial prompts to expose weaknesses, inconsistencies, and unexpected behavior.
I assess model outputs for factuality, reasoning, instruction following, contextual understanding, relevance, and quality. I create evaluation rubrics, identify recurring failure modes, rank alternative responses, and provide structured feedback for improvement.
My work also includes AI audio evaluation at CrowdGen and multilingual search quality evaluation at Welocalize/Welo Data. I bring a Quality Assurance background from McCain Foods, along with technical knowledge of full-stack development, web technologies, programming fundamentals, and data analysis.

