At Metis, I shipped a production RAG pipeline covering ingestion, retrieval, prompt construction, and the OpenAI API layer. I traced poor answer quality to fixed-size chunking that separated financial tables from their headers and footnotes, then rebuilt chunking around document structure.
I also implemented hybrid Elasticsearch and dense-vector retrieval with reranking, and built a provider abstraction with token budgeting, retries, and failure handling. I defined accuracy KPIs and reported where outputs were ready for production and where human review remained necessary.
On my Voice-AI Turn-Detection Benchmark project, I evaluated six strategies across 1,152 labelled speech turns, measuring latency alongside interruption rate. I caught and corrected a latency measurement bug, retracted two claims after paired reanalysis, and built an async testable pipeline.

