At Instead, I own output quality for a production tax-research platform, where a wrong citation can create a compliance risk. I built the evaluation layer that turns failure reviews into product fixes, and the golden evaluation set holds about 95% citation accuracy.
I also built a self-maintaining corpus of 270K+ primary-law records and designed Keystone, an internal benchmark, alongside Project Evals, a human-expert ground-truth program targeting 2,000 reviewed datasets. These foundations help the team measure retrieval and answer quality and guide model decisions with evidence.
Before Instead, I shipped a background verification module at Keka HR that opened the company’s first international market, onboarding 28 US enterprise clients and contributing $2.7M in MRR. I’ve also published CrossSource and mirror-eval, open-source evaluation harnesses for citation reliability and AI search retrieval.

