At Paradigm Networks, Inc., I owned the full pipeline for a RAG factuality detector, from curating training samples and fine-tuning IBM Granite Guardian 3.1 to production deployment. It achieved 91.7% accuracy and 86.7% F1, outperforming the listed GPT-5 and Claude 3 Sonnet results.
I also built an agentic evaluation platform that scores AI-agent runs across nine metrics, and co-developed and served a hallucination detector that reached 98.1% F1. In my MS Data Science research, I built a retrieval pipeline for small LLMs that raised held-out accuracy from 75.0% to 85.4%.

