At Outlier AI, I audit and rank complex LLM outputs against rubrics for reasoning, safety, and goal compliance. I document agent failure modes and design error-classification taxonomies for model alignment and benchmarking.
At the University of Alaska, I develop reproducible Python pipelines to simulate multi-turn, goal-oriented web browsing. I compare synthetic clickstream patterns with real-world user interaction traces.
I also conduct significance testing and ablation studies to examine how prompt variations and context constraints affect model behavior.
Earlier, as a Research & Academic Content Assistant at University of Alaska Anchorage, I synthesized qualitative evaluation reports and checked research data against original empirical sources. My Ph.D. work focuses on simulated user exploration and behavioral calibration in autonomous web agents.

