At Handshake AI, I evaluated and ranked large language model responses for accuracy, helpfulness, tone, and adherence to task guidelines. I also reviewed AI-generated code, flagged logic errors, and suggested corrected implementations.
At Mercor, I fact-checked AI-generated content and verified sources before dataset submission. I compared model responses and provided structured feedback for fine-tuning, including in coding-focused evaluations.
At Outlier, I rated AI responses for helpfulness, coherence, and safety in RLHF training pipelines. I wrote prompts and reference answers, and reported inconsistencies, hallucinations, and factual errors.

