At HandshakeAI, I designed and validated deterministic, container-based AI evaluation benchmarks with reproducible scoring. I defined hidden reference data, output schemas, numerical tolerances, and automated test harnesses.
I also built data-preparation pipelines that turned raw time-series records into analysis-ready datasets, and developed scoring logic using weighted error metrics and prediction-interval validation.
At Mercor, I conducted LLM training and evaluation through prompt generation and rubric grading. I assessed model responses for factual accuracy and logical consistency, including technical questions and AI-generated code.
My projects include AetherMind, a deployed full-stack AI web platform, and an AI resume analyzer. I also built web, quiz-management, and database-driven projects using tools and technologies including React, Python, Java, and SQL.

