At TELUS Digital AI, I evaluated large language model responses for factuality, safety, relevance, and instruction following. I analyzed failure patterns and translated my findings into structured recommendations for AI training and model improvement.
At Sama and RWS, I supported AI training workflows through data annotation, validation, and quality assessment. My research projects explored synthetic-data evaluation, LLM-as-a-Judge methods, and repeatable workflows for assessing model behavior.

