At Outlier / Scale AI, I evaluate LLM outputs against technical rubrics for code correctness, mathematical logic, instruction following, and safety. I built RLHF, preference-ranking, and red-teaming workflows in Python that reduced model hallucinations by 32%.
I also created automation scripts and text-parsing tools to support error analysis, document failure patterns, and prioritize fine-tuning for foundation-model releases. Across evaluation cycles, I identified edge-case failures and contributed to improved model accuracy on MMLU-Pro benchmark tasks.
At Braze, I implemented and maintained backend RESTful microservices and data integration pipelines using Python, SQL, and AWS. I automated QA workflows and backend integration tests, reducing the system defect rate by 35%.
At Appen, I evaluated search relevance and API responses, audited web application outputs, and identified bias, logic errors, and policy violations in generative AI training data. Earlier, at Nairobi Data Labs, I built SQL pipelines to validate data assumptions for modeling projects.

