I've built LLM training and evaluation pipelines at Turing.com for Apple, generating and validating training and labeling data with OpenAI models. I also designed rubric-based and LLM-as-a-judge workflows that improved model-quality tracking and iteration.
My recent work includes Terminal-Bench evaluation reliability and reporting, plus a multi-agent system in LangChain and LangGraph with planner, researcher, coder, and reviewer agents. I've worked across OpenAI, Claude Sonnet, Gemini, and Qwen3-Coder for coding and agent workflows.
At Cognostics AG, I built backend services, recommendation workflows, and APIs for a digital health app tracking diet, exercise, and sleep.
At Chegg, I developed personalization pipelines using billion-plus-row datasets and Kafka, Kinesis, and Redshift, and helped build an online tutoring platform with audio/video calling, screen sharing, and collaborative code editing. Earlier, I built a SaaS Gmail extension at Grexit and automation features for a manufacturing client at Nagarro Software.
