At ICON plc, I built the AI evaluation pipeline that each release of an AI-assisted clinical trial monitoring platform must pass. It combines RAGAS metrics, LLM-as-judge scoring, and a golden dataset created with clinical experts, alongside checks for hallucinations, prompt injection, PII, and toxicity.
At Trace Labs / OriginTrail, I set up RAG evaluation for a knowledge-graph-based AI assistant and wired safety checks into CI/CD as release gates. Earlier, at TeleTrader, I built and maintained Playwright regression tests for the Baha Webstation trading platform.

