At Maggu, I reworked the LLM judge for medicine recommendations, raising accuracy from 62% to 94% with adversarial agents and a human-made evaluation dataset. I also tracked operational costs, latency, accuracy, and routing in Databricks MLflow.
At Maggu, I cut time to first token from 1.5 seconds to 0.7 seconds by revising cache and RAG strategies. I provisioned reproducible Databricks pipelines with Terraform and reduced compute spend by profiling PySpark jobs.
At Capgemini, I wrote an evaluation suite for agent workflows that caught four regressions before they reached the client. I also ran workshops on agent design and LLM-as-a-judge for a newly formed team and two managers from Vivo.
Across my work at Acaso and Di2win, I built RAG and model-training pipelines, automated document ingestion, and shipped generative AI proofs of concept. I bring a computer vision research background from UFPE and have worked across data pipelines, backend systems, and production ML infrastructure.

