At Alignerr (Labelbox) and Mercor, I author benchmark tasks for AI coding agents and evaluate agentic coding transcripts and model-generated fixes against structured rubrics.
I also review LLM-generated solutions to ML problems, write justifications about correctness and feasibility, and maintain Python build scripts across open-source repositories for training data pipelines.
At Accenture, I built FastAPI services and document ingestion pipelines that supported RAG and LLM evaluation. I also developed a GenAI Configurator for comparing OpenAI, Meta Llama, and Anthropic Claude models.
At Accenture, I led GitHub Copilot performance evaluations and supported its rollout, and deployed GenAI applications with Docker and Kubernetes. Earlier, at Cleverground, I built Django REST Framework backend features for a learning management system and wrote Python and Selenium test automation.

