I've built research evaluation systems for large language models, including SAFI, which benchmarks frontier LLM capabilities across all 35 O*NET workforce skills. I designed and ran 263 benchmark tasks across LLaMA 3.3 70B, Mistral Large, Qwen 2.5 72B, and Gemini 2.5 Flash, producing 1,052 scored evaluations and an AI Impact Matrix.
My research also investigates fairness in AI-assisted assessment. I developed a controlled framework using 180 student responses and 480 automated grading evaluations, finding that writing style can significantly skew LLM grading even with explicit debiasing prompts.
I also build practical software, from Cortexa, a fully client-side Chrome extension for detecting context drift in ChatGPT conversations, to a full-stack membership platform for the Association of Computer Engineering Students. I work with Python, JavaScript, Node.js, Express.js, MySQL, and machine-learning data tools.
