At the University of Maryland, I built a framework that evaluates LLM agents against known ground truth from a synthetic environment. I also benchmarked a scripted baseline against PC and GES causal structure-learning algorithms.
At Nativebyte, I built and deployed Python ML pipelines behind XGBoost and regression models in a live product. In my projects, I caught data leakage in phishing-email classification and fine-tuned T5 for medical text summarization.

